Grapharizer: A Graph-Based Technique for Extractive Multi-Document Summarization

被引:6
|
作者
Jalil, Zakia [1 ]
Nasir, Muhammad [2 ]
Alazab, Moutaz [3 ]
Nasir, Jamal [4 ]
Amjad, Tehmina [1 ]
Alqammaz, Abdullah [5 ]
机构
[1] Int Islamic Univ, Dept Comp Sci, Islamabad 44000, Pakistan
[2] Int Islamic Univ, Dept Software Engn, Islamabad 44000, Pakistan
[3] Al Balqa Appl Univ, Fac Artificial Intelligence, Dept Intelligent Syst, Salt 19117, Jordan
[4] Univ Galway, Sch Comp Sci, Galway H91TK33, Ireland
[5] Zarqa Univ, Coll Informat Technol, Dept Cyber Secur, Zarqa 13110, Jordan
关键词
big data; automatic text summarization; extractive multi-document summarization; graph theory; machine learning; anaphora; cataphora; pronoun resolution; grammaticality; topic modeling; ChatGPT; TEXT; SEARCH;
D O I
10.3390/electronics12081895
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Featured Application A graph-based technique tested on a benchmark dataset and augmented by machine learning techniques to provide a concise, informative, and grammatically correct summary. In the age of big data, there is increasing growth of data on the Internet. It becomes frustrating for users to locate the desired data. Therefore, text summarization emerges as a solution to this problem. It summarizes and presents the users with the gist of the provided documents. However, summarizer systems face challenges, such as poor grammaticality, missing important information, and redundancy, particularly in multi-document summarization. This study involves the development of a graph-based extractive generic MDS technique, named Grapharizer (GRAPH-based summARIZER), focusing on resolving these challenges. Grapharizer addresses the grammaticality problems of the summary using lemmatization during pre-processing. Furthermore, synonym mapping, multi-word expression mapping, and anaphora and cataphora resolution, contribute positively to improving the grammaticality of the generated summary. Challenges, such as redundancy and proper coverage of all topics, are dealt with to achieve informativity and representativeness. Grapharizer is a novel approach which can also be used in combination with different machine learning models. The system was tested on DUC 2004 and Recent News Article datasets against various state-of-the-art techniques. Use of Grapharizer with machine learning increased accuracy by up to 23.05% compared with different baseline techniques on ROUGE scores. Expert evaluation of the proposed system indicated the accuracy to be more than 55%.
引用
收藏
页数:26
相关论文
共 50 条
  • [1] Unsupervised Graph-Based Tibetan Multi-Document Summarization
    Yan, Xiaodong
    Wang, Yiqin
    Wei Song
    Zhao, Xiaobing
    Run, A.
    Yang Yanxing
    CMC-COMPUTERS MATERIALS & CONTINUA, 2022, 73 (01): : 1769 - 1781
  • [2] Multi-document extractive summarization using semantic graph
    del Camino Valle, Oleyda
    Simon-Cuevas, Alfredo
    Valladares-Valdes, Eduardo
    Olivas, Jose A.
    Romero, Francisco P.
    PROCESAMIENTO DEL LENGUAJE NATURAL, 2019, (63): : 103 - 110
  • [3] Extractive multi-document text summarization based on graph independent sets
    Uckan, Taner
    Karci, Ali
    EGYPTIAN INFORMATICS JOURNAL, 2020, 21 (03) : 145 - 157
  • [4] An Extractive Multi-Document Summarization Technique Based on Fuzzy Logic approach
    Tsoumou, Evrard Stency Larys
    Yang, Shichong
    Lai, Linjing
    Varus, Mbembo Loundou
    2016 INTERNATIONAL CONFERENCE ON NETWORK AND INFORMATION SYSTEMS FOR COMPUTERS (ICNISC), 2016, : 346 - 351
  • [5] Enhancing Multi-Document Summarization with Cross-Document Graph-based Information Extraction
    Zhang, Zixuan
    Elfardy, Heba
    Dreyer, Markus
    Small, Kevin
    Ji, Heng
    Bansal, Mohit
    17TH CONFERENCE OF THE EUROPEAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, EACL 2023, 2023, : 1696 - 1707
  • [6] Graph-based extractive text summarization based on single document
    Avaneesh Kumar Yadav
    Rama Shankar Ranvijay
    Ashish Kumar Yadav
    Multimedia Tools and Applications, 2024, 83 : 18987 - 19013
  • [7] Graph-based extractive text summarization based on single document
    Yadav, Avaneesh Kumar
    Ranvijay, Rama Shankar
    Yadav, Rama Shankar
    Maurya, Ashish Kumar
    MULTIMEDIA TOOLS AND APPLICATIONS, 2024, 83 (07) : 18987 - 19013
  • [8] Graph-Based Query-Focused Multi-document Summarization Using Improved Affinity Graph
    Hu, Po
    He, Jiacong
    Zhang, Yong
    KNOWLEDGE SCIENCE, ENGINEERING AND MANAGEMENT, KSEM 2015, 2015, 9403 : 336 - 347
  • [9] Multi-document extractive text summarization based on firefly algorithm
    Tomer, Minakshi
    Kumar, Manoj
    JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES, 2022, 34 (08) : 6057 - 6065
  • [10] Graph-Based Multi-Modality Learning for Topic-Focused Multi-Document Summarization
    Wan, Xiaojun
    Xiao, Jianguo
    21ST INTERNATIONAL JOINT CONFERENCE ON ARTIFICIAL INTELLIGENCE (IJCAI-09), PROCEEDINGS, 2009, : 1586 - 1591