The task of identifying Chinese named entities of Chinese poetry and wine culture is a key step in the construction of a knowledge graph and a question and answer system.Aimed at the characteristics of Chinese poetry ...The task of identifying Chinese named entities of Chinese poetry and wine culture is a key step in the construction of a knowledge graph and a question and answer system.Aimed at the characteristics of Chinese poetry and wine culture entities with different lengths and high training cost of named entity recognition models at the present stage,this study proposes a lite BERT+bi-directional long short-term memory+attentional mechanisms+conditional random field(ALBERT+BILSTM+Att+CRF).The method first obtains the characterlevel semantic information by ALBERT module,then extracts its high-dimensional features by BILSTM module,weights the original word vector and the learned text vector by attention layer,and finally predicts the true label in CRF module(including five types:poem title,author,time,genre,and category).Through experiments on data sets related to Chinese poetry and wine culture,the results show that the method is more effective than existing mainstream models and can efficiently extract important entity information in Chinese poetry and wine culture,which is an effective method for the identification of named entities of varying lengths of poetry.展开更多
基金the Sichuan Science and Technology Program of China(No.2021YFG0055)the Zigong Science and Technology Program of China(No.2019YYJC15)+1 种基金the Nature Science Foundation of Sichuan University of Science&Engineering(No.2020RC32)the 2022 Graduate Innovation Fund Project of Sichuan University of Science&Engineering(No.Y2022168)。
文摘The task of identifying Chinese named entities of Chinese poetry and wine culture is a key step in the construction of a knowledge graph and a question and answer system.Aimed at the characteristics of Chinese poetry and wine culture entities with different lengths and high training cost of named entity recognition models at the present stage,this study proposes a lite BERT+bi-directional long short-term memory+attentional mechanisms+conditional random field(ALBERT+BILSTM+Att+CRF).The method first obtains the characterlevel semantic information by ALBERT module,then extracts its high-dimensional features by BILSTM module,weights the original word vector and the learned text vector by attention layer,and finally predicts the true label in CRF module(including five types:poem title,author,time,genre,and category).Through experiments on data sets related to Chinese poetry and wine culture,the results show that the method is more effective than existing mainstream models and can efficiently extract important entity information in Chinese poetry and wine culture,which is an effective method for the identification of named entities of varying lengths of poetry.