Cargando…
Deep Modular Bilinear Attention Network for Visual Question Answering
VQA (Visual Question Answering) is a multi-model task. Given a picture and a question related to the image, it will determine the correct answer. The attention mechanism has become a de facto component of almost all VQA models. Most recent VQA approaches use dot-product to calculate the intra-modali...
Autores principales: | Yan, Feng, Silamu, Wushouer, Li, Yanbing |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
MDPI
2022
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8838230/ https://www.ncbi.nlm.nih.gov/pubmed/35161790 http://dx.doi.org/10.3390/s22031045 |
Ejemplares similares
-
An Effective Dense Co-Attention Networks for Visual Question Answering
por: He, Shirong, et al.
Publicado: (2020) -
Multi-Modal Explicit Sparse Attention Networks for Visual Question Answering
por: Guo, Zihan, et al.
Publicado: (2020) -
Adversarial Learning with Bidirectional Attention for Visual Question Answering
por: Li, Qifeng, et al.
Publicado: (2021) -
Focal cross transformer: multi-view brain tumor segmentation model based on cross window and focal self-attention
por: Zongren, Li, et al.
Publicado: (2023) -
Research on visual question answering based on dynamic memory network model of multiple attention mechanisms
por: Miao, Yalin, et al.
Publicado: (2022)