期刊文献+
共找到1篇文章
< 1 >
每页显示 20 50 100
Dual modality prompt learning for visual question-grounded answering in robotic surgery 被引量:2
1
作者 Yue Zhang Wanshu Fan +3 位作者 peixi peng Xin Yang Dongsheng Zhou Xiaopeng Wei 《Visual Computing for Industry,Biomedicine,and Art》 2024年第1期316-328,共13页
With recent advancements in robotic surgery,notable strides have been made in visual question answering(VQA).Existing VQA systems typically generate textual answers to questions but fail to indicate the location of th... With recent advancements in robotic surgery,notable strides have been made in visual question answering(VQA).Existing VQA systems typically generate textual answers to questions but fail to indicate the location of the relevant content within the image.This limitation restricts the interpretative capacity of the VQA models and their abil-ity to explore specific image regions.To address this issue,this study proposes a grounded VQA model for robotic surgery,capable of localizing a specific region during answer prediction.Drawing inspiration from prompt learning in language models,a dual-modality prompt model was developed to enhance precise multimodal information interactions.Specifically,two complementary prompters were introduced to effectively integrate visual and textual prompts into the encoding process of the model.A visual complementary prompter merges visual prompt knowl-edge with visual information features to guide accurate localization.The textual complementary prompter aligns vis-ual information with textual prompt knowledge and textual information,guiding textual information towards a more accurate inference of the answer.Additionally,a multiple iterative fusion strategy was adopted for comprehensive answer reasoning,to ensure high-quality generation of textual and grounded answers.The experimental results vali-date the effectiveness of the model,demonstrating its superiority over existing methods on the EndoVis-18 and End-oVis-17 datasets. 展开更多
关键词 Prompt learning Visual prompt Textual prompt Grounding-answering Visual question answering
在线阅读 下载PDF
上一页 1 下一页 到第
使用帮助 返回顶部