Grapharizer: A Graph-Based Technique for Extractive Multi-Document Summarization

被引:6
|
作者
Jalil, Zakia [1 ]
Nasir, Muhammad [2 ]
Alazab, Moutaz [3 ]
Nasir, Jamal [4 ]
Amjad, Tehmina [1 ]
Alqammaz, Abdullah [5 ]
机构
[1] Int Islamic Univ, Dept Comp Sci, Islamabad 44000, Pakistan
[2] Int Islamic Univ, Dept Software Engn, Islamabad 44000, Pakistan
[3] Al Balqa Appl Univ, Fac Artificial Intelligence, Dept Intelligent Syst, Salt 19117, Jordan
[4] Univ Galway, Sch Comp Sci, Galway H91TK33, Ireland
[5] Zarqa Univ, Coll Informat Technol, Dept Cyber Secur, Zarqa 13110, Jordan
关键词
big data; automatic text summarization; extractive multi-document summarization; graph theory; machine learning; anaphora; cataphora; pronoun resolution; grammaticality; topic modeling; ChatGPT; TEXT; SEARCH;
D O I
10.3390/electronics12081895
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Featured Application A graph-based technique tested on a benchmark dataset and augmented by machine learning techniques to provide a concise, informative, and grammatically correct summary. In the age of big data, there is increasing growth of data on the Internet. It becomes frustrating for users to locate the desired data. Therefore, text summarization emerges as a solution to this problem. It summarizes and presents the users with the gist of the provided documents. However, summarizer systems face challenges, such as poor grammaticality, missing important information, and redundancy, particularly in multi-document summarization. This study involves the development of a graph-based extractive generic MDS technique, named Grapharizer (GRAPH-based summARIZER), focusing on resolving these challenges. Grapharizer addresses the grammaticality problems of the summary using lemmatization during pre-processing. Furthermore, synonym mapping, multi-word expression mapping, and anaphora and cataphora resolution, contribute positively to improving the grammaticality of the generated summary. Challenges, such as redundancy and proper coverage of all topics, are dealt with to achieve informativity and representativeness. Grapharizer is a novel approach which can also be used in combination with different machine learning models. The system was tested on DUC 2004 and Recent News Article datasets against various state-of-the-art techniques. Use of Grapharizer with machine learning increased accuracy by up to 23.05% compared with different baseline techniques on ROUGE scores. Expert evaluation of the proposed system indicated the accuracy to be more than 55%.
引用
收藏
页数:26
相关论文
共 50 条
  • [21] SRRank: Leveraging Semantic Roles for Extractive Multi-Document Summarization
    Yan, Su
    Wan, Xiaojun
    IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2014, 22 (12) : 2048 - 2058
  • [22] Survey on Extractive Text Summarization Methods with Multi-Document Datasets
    Varalakshmi, P. N. K.
    Kallimani, Jagadish S.
    2018 INTERNATIONAL CONFERENCE ON ADVANCES IN COMPUTING, COMMUNICATIONS AND INFORMATICS (ICACCI), 2018, : 2113 - 2119
  • [23] A Preliminary Exploration of Extractive Multi-Document Summarization in Hyperbolic Space
    Song, Mingyang
    Feng, Yi
    Jing, Liping
    PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON INFORMATION AND KNOWLEDGE MANAGEMENT, CIKM 2022, 2022, : 4505 - 4509
  • [24] A document-sensitive graph model for multi-document summarization
    Furu Wei
    Wenjie Li
    Qin Lu
    Yanxiang He
    Knowledge and Information Systems, 2010, 22 : 245 - 259
  • [25] Multi-document Extractive Summarization Using Window-based Sentence Representation
    Zhang, Yong
    Er, Meng Joo
    Zhao, Rui
    2015 IEEE SYMPOSIUM SERIES ON COMPUTATIONAL INTELLIGENCE (IEEE SSCI), 2015, : 404 - 410
  • [26] MSCSO: Extractive Multi-document Summarization Based on a New Criterion of Sentences Overlapping
    Khaleghi, Zeynab
    Fakhredanesh, Mohammad
    Hourali, Maryam
    IRANIAN JOURNAL OF SCIENCE AND TECHNOLOGY-TRANSACTIONS OF ELECTRICAL ENGINEERING, 2021, 45 (01) : 195 - 205
  • [27] A hybrid model for sentence ordering in extractive multi-document summarization
    Liu, Dexi
    Zhang, Zengchang
    He, Yanxiang
    Ji, Donghong
    INFORMATION RETRIEVAL TECHNOLOGY, PROCEEDINGS, 2006, 4182 : 588 - 592
  • [28] Extractive Multi-Document Summarization: A Review of Progress in the Last Decade
    Jalil, Zakia
    Nasir, Jamal Abdul
    Nasir, Muhammad
    IEEE ACCESS, 2021, 9 : 130928 - 130946
  • [29] MSCSO: Extractive Multi-document Summarization Based on a New Criterion of Sentences Overlapping
    Zeynab Khaleghi
    Mohammad Fakhredanesh
    Maryam Hourali
    Iranian Journal of Science and Technology, Transactions of Electrical Engineering, 2021, 45 : 195 - 205
  • [30] Extractive multi-document summarization using population-based multicriteria optimization
    John, Ansamma
    Premjith, P. S.
    Wilscy, M.
    EXPERT SYSTEMS WITH APPLICATIONS, 2017, 86 : 385 - 397