摘要
为了探索图像场景理解所需要的视觉区域间关系的建模与推理,提出视觉关系推理模块.该模块基于图像中不同的语义和空间上下文信息,对相关视觉对象间的关系模式进行动态编码,并推断出与当前生成的关系词最相关的语义特征输出.通过引入上下文门控机制,以根据不同类型的单词动态地权衡视觉注意力模块和视觉关系推理模块的贡献.实验结果表明,对比以往基于注意力机制的图像描述方法,基于视觉关系推理与上下文门控机制的图像描述方法更好;所提模块可以动态建模和推理不同类型生成单词的最相关特征,对输入图像中物体关系的描述更加准确.
A visual relationship reasoning module was proposed in order to explore the modeling and reasoning of the relationship between visual regions needed for image scene understanding.The relationship patterns between the two related visual objects were encoded dynamically based on different semantic and spatial context information,and the most relevant feature output of the currently generated relationship words was inferred by using the module.In addition,the contributions between the visual attention module and the visual relational reasoning module were controlled dynamically according to the different types of words by introducing the context gate mechanism.Experimental results show that the method has better performance than other state-of-the-art methods based on attention mechanism.By using the module a model is established dynamically,the most relevant features of different types for the generated words are inferred,and the quality of image caption is improved.
作者
陈巧红
裴皓磊
孙麒
CHEN Qiao-hong;PEI Hao-lei;SUN Qi(School of Informatics Science and Technology,Zhejiang Sci-Tech University,Hangzhou 310018,China)
出处
《浙江大学学报(工学版)》
EI
CAS
CSCD
北大核心
2022年第3期542-549,共8页
Journal of Zhejiang University:Engineering Science