Loading…

Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering

Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both...

Full description

Saved in:

Bibliographic Details
Published in:	IEEE transaction on neural networks and learning systems 2018-12, Vol.29 (12), p.5947-5959
Main Authors:	Yu, Zhou, Yu, Jun, Xiang, Chenchao, Fan, Jianping, Tao, Dacheng
Format:	Article
Language:	English
Subjects:	Artificial neural networks Coattention learning Computational modeling Correlation deep learning Divergence Feature extraction Knowledge discovery Mathematical models multimodal feature fusion Natural languages Neural networks Representations State of the art Task analysis visual question answering (VQA) Visualization
Citations:	Items that this one cites Items that cite this one
Online Access:	Get full text
Tags:	Add Tag No Tags, Be the first to tag this record!

cited_by	cdi_FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413
cites	cdi_FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413
container_end_page	5959
container_issue	12
container_start_page	5947
container_title	IEEE transaction on neural networks and learning systems
container_volume	29
creator	Yu, Zhou Yu, Jun Xiang, Chenchao Fan, Jianping Tao, Dacheng
description	Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.
doi_str_mv	10.1109/TNNLS.2018.2817340
format	article
fullrecord	<record><control><sourceid>proquest_ieee_</sourceid><recordid>TN_cdi_ieee_primary_8334194</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><ieee_id>8334194</ieee_id><sourcerecordid>2068342533</sourcerecordid><originalsourceid>FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413</originalsourceid><addsrcrecordid>eNpdkVtLxDAQhYMoKuofUJCCL750zSRpm_im4g3WG17wLaTNrEa6jSZbRH-9WXfdB_OSMPOdw0wOIdtABwBUHTxcXw_vB4yCHDAJFRd0iawzKFnOuJTLi3f1vEa2Ynyj6ZS0KIVaJWtMKcWlqNZJc4xfvrPZsWtdhyYcZufYYTCt-0abXfXtxI29NW12ZpqJD7_VC_fymt8EiyG79T7pXrKRD9mTi30C73qME-e77KiLnxhSd5OsjEwbcWt-b5DHs9OHk4t8eHN-eXI0zBsB1SRvZKUYNIyjNVaUBY6YqKtaSckbKARgQRXnYGvJS4pMgaCsVFAbsDYtA3yD7M9834P_mE6hxy422LamQ99HzWgpuWAF5wnd-4e--T50aTrNgFeFlBUrE8VmVBN8jAFH-j24sQlfGqiepqB_U9DTFPQ8hSTanVv39RjtQvL35wnYmQEOERdtybkAJfgPLtGKVQ</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2137588726</pqid></control><display><type>article</type><title>Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering</title><source>IEEE Electronic Library (IEL) Journals</source><creator>Yu, Zhou ; Yu, Jun ; Xiang, Chenchao ; Fan, Jianping ; Tao, Dacheng</creator><creatorcontrib>Yu, Zhou ; Yu, Jun ; Xiang, Chenchao ; Fan, Jianping ; Tao, Dacheng</creatorcontrib><description>Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.</description><identifier>ISSN: 2162-237X</identifier><identifier>EISSN: 2162-2388</identifier><identifier>DOI: 10.1109/TNNLS.2018.2817340</identifier><identifier>PMID: 29993847</identifier><identifier>CODEN: ITNNAL</identifier><language>eng</language><publisher>United States: IEEE</publisher><subject>Artificial neural networks ; Coattention learning ; Computational modeling ; Correlation ; deep learning ; Divergence ; Feature extraction ; Knowledge discovery ; Mathematical models ; multimodal feature fusion ; Natural languages ; Neural networks ; Representations ; State of the art ; Task analysis ; visual question answering (VQA) ; Visualization</subject><ispartof>IEEE transaction on neural networks and learning systems, 2018-12, Vol.29 (12), p.5947-5959</ispartof><rights>Copyright The Institute of Electrical and Electronics Engineers, Inc. (IEEE) 2018</rights><woscitedreferencessubscribed>false</woscitedreferencessubscribed><citedby>FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413</citedby><cites>FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413</cites><orcidid>0000-0001-7225-5449 ; 0000-0001-8407-1137 ; 0000-0002-4923-0910 ; 0000-0003-1922-7283</orcidid></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><linktohtml>$$Uhttps://ieeexplore.ieee.org/document/8334194$$EHTML$$P50$$Gieee$$H</linktohtml><link.rule.ids>314,780,784,27924,27925,54796</link.rule.ids><backlink>$$Uhttps://www.ncbi.nlm.nih.gov/pubmed/29993847$$D View this record in MEDLINE/PubMed$$Hfree_for_read</backlink></links><search><creatorcontrib>Yu, Zhou</creatorcontrib><creatorcontrib>Yu, Jun</creatorcontrib><creatorcontrib>Xiang, Chenchao</creatorcontrib><creatorcontrib>Fan, Jianping</creatorcontrib><creatorcontrib>Tao, Dacheng</creatorcontrib><title>Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering</title><title>IEEE transaction on neural networks and learning systems</title><addtitle>TNNLS</addtitle><addtitle>IEEE Trans Neural Netw Learn Syst</addtitle><description>Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.</description><subject>Artificial neural networks</subject><subject>Coattention learning</subject><subject>Computational modeling</subject><subject>Correlation</subject><subject>deep learning</subject><subject>Divergence</subject><subject>Feature extraction</subject><subject>Knowledge discovery</subject><subject>Mathematical models</subject><subject>multimodal feature fusion</subject><subject>Natural languages</subject><subject>Neural networks</subject><subject>Representations</subject><subject>State of the art</subject><subject>Task analysis</subject><subject>visual question answering (VQA)</subject><subject>Visualization</subject><issn>2162-237X</issn><issn>2162-2388</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2018</creationdate><recordtype>article</recordtype><recordid>eNpdkVtLxDAQhYMoKuofUJCCL750zSRpm_im4g3WG17wLaTNrEa6jSZbRH-9WXfdB_OSMPOdw0wOIdtABwBUHTxcXw_vB4yCHDAJFRd0iawzKFnOuJTLi3f1vEa2Ynyj6ZS0KIVaJWtMKcWlqNZJc4xfvrPZsWtdhyYcZufYYTCt-0abXfXtxI29NW12ZpqJD7_VC_fymt8EiyG79T7pXrKRD9mTi30C73qME-e77KiLnxhSd5OsjEwbcWt-b5DHs9OHk4t8eHN-eXI0zBsB1SRvZKUYNIyjNVaUBY6YqKtaSckbKARgQRXnYGvJS4pMgaCsVFAbsDYtA3yD7M9834P_mE6hxy422LamQ99HzWgpuWAF5wnd-4e--T50aTrNgFeFlBUrE8VmVBN8jAFH-j24sQlfGqiepqB_U9DTFPQ8hSTanVv39RjtQvL35wnYmQEOERdtybkAJfgPLtGKVQ</recordid><startdate>20181201</startdate><enddate>20181201</enddate><creator>Yu, Zhou</creator><creator>Yu, Jun</creator><creator>Xiang, Chenchao</creator><creator>Fan, Jianping</creator><creator>Tao, Dacheng</creator><general>IEEE</general><general>The Institute of Electrical and Electronics Engineers, Inc. (IEEE)</general><scope>97E</scope><scope>RIA</scope><scope>RIE</scope><scope>NPM</scope><scope>AAYXX</scope><scope>CITATION</scope><scope>7QF</scope><scope>7QO</scope><scope>7QP</scope><scope>7QQ</scope><scope>7QR</scope><scope>7SC</scope><scope>7SE</scope><scope>7SP</scope><scope>7SR</scope><scope>7TA</scope><scope>7TB</scope><scope>7TK</scope><scope>7U5</scope><scope>8BQ</scope><scope>8FD</scope><scope>F28</scope><scope>FR3</scope><scope>H8D</scope><scope>JG9</scope><scope>JQ2</scope><scope>KR7</scope><scope>L7M</scope><scope>L~C</scope><scope>L~D</scope><scope>P64</scope><scope>7X8</scope><orcidid>https://orcid.org/0000-0001-7225-5449</orcidid><orcidid>https://orcid.org/0000-0001-8407-1137</orcidid><orcidid>https://orcid.org/0000-0002-4923-0910</orcidid><orcidid>https://orcid.org/0000-0003-1922-7283</orcidid></search><sort><creationdate>20181201</creationdate><title>Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering</title><author>Yu, Zhou ; Yu, Jun ; Xiang, Chenchao ; Fan, Jianping ; Tao, Dacheng</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2018</creationdate><topic>Artificial neural networks</topic><topic>Coattention learning</topic><topic>Computational modeling</topic><topic>Correlation</topic><topic>deep learning</topic><topic>Divergence</topic><topic>Feature extraction</topic><topic>Knowledge discovery</topic><topic>Mathematical models</topic><topic>multimodal feature fusion</topic><topic>Natural languages</topic><topic>Neural networks</topic><topic>Representations</topic><topic>State of the art</topic><topic>Task analysis</topic><topic>visual question answering (VQA)</topic><topic>Visualization</topic><toplevel>online_resources</toplevel><creatorcontrib>Yu, Zhou</creatorcontrib><creatorcontrib>Yu, Jun</creatorcontrib><creatorcontrib>Xiang, Chenchao</creatorcontrib><creatorcontrib>Fan, Jianping</creatorcontrib><creatorcontrib>Tao, Dacheng</creatorcontrib><collection>IEEE All-Society Periodicals Package (ASPP) 2005-present</collection><collection>IEEE All-Society Periodicals Package (ASPP) 1998-Present</collection><collection>IEEE Xplore</collection><collection>PubMed</collection><collection>CrossRef</collection><collection>Aluminium Industry Abstracts</collection><collection>Biotechnology Research Abstracts</collection><collection>Calcium & Calcified Tissue Abstracts</collection><collection>Ceramic Abstracts</collection><collection>Chemoreception Abstracts</collection><collection>Computer and Information Systems Abstracts</collection><collection>Corrosion Abstracts</collection><collection>Electronics & Communications Abstracts</collection><collection>Engineered Materials Abstracts</collection><collection>Materials Business File</collection><collection>Mechanical & Transportation Engineering Abstracts</collection><collection>Neurosciences Abstracts</collection><collection>Solid State and Superconductivity Abstracts</collection><collection>METADEX</collection><collection>Technology Research Database</collection><collection>ANTE: Abstracts in New Technology & Engineering</collection><collection>Engineering Research Database</collection><collection>Aerospace Database</collection><collection>Materials Research Database</collection><collection>ProQuest Computer Science Collection</collection><collection>Civil Engineering Abstracts</collection><collection>Advanced Technologies Database with Aerospace</collection><collection>Computer and Information Systems Abstracts Academic</collection><collection>Computer and Information Systems Abstracts Professional</collection><collection>Biotechnology and BioEngineering Abstracts</collection><collection>MEDLINE - Academic</collection><jtitle>IEEE transaction on neural networks and learning systems</jtitle></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Yu, Zhou</au><au>Yu, Jun</au><au>Xiang, Chenchao</au><au>Fan, Jianping</au><au>Tao, Dacheng</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering</atitle><jtitle>IEEE transaction on neural networks and learning systems</jtitle><stitle>TNNLS</stitle><addtitle>IEEE Trans Neural Netw Learn Syst</addtitle><date>2018-12-01</date><risdate>2018</risdate><volume>29</volume><issue>12</issue><spage>5947</spage><epage>5959</epage><pages>5947-5959</pages><issn>2162-237X</issn><eissn>2162-2388</eissn><coden>ITNNAL</coden><abstract>Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.</abstract><cop>United States</cop><pub>IEEE</pub><pmid>29993847</pmid><doi>10.1109/TNNLS.2018.2817340</doi><tpages>13</tpages><orcidid>https://orcid.org/0000-0001-7225-5449</orcidid><orcidid>https://orcid.org/0000-0001-8407-1137</orcidid><orcidid>https://orcid.org/0000-0002-4923-0910</orcidid><orcidid>https://orcid.org/0000-0003-1922-7283</orcidid></addata></record>
fulltext	fulltext
identifier	ISSN: 2162-237X
ispartof	IEEE transaction on neural networks and learning systems, 2018-12, Vol.29 (12), p.5947-5959
issn	2162-237X 2162-2388
language	eng
recordid	cdi_ieee_primary_8334194
source	IEEE Electronic Library (IEL) Journals
subjects	Artificial neural networks Coattention learning Computational modeling Correlation deep learning Divergence Feature extraction Knowledge discovery Mathematical models multimodal feature fusion Natural languages Neural networks Representations State of the art Task analysis visual question answering (VQA) Visualization
title	Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering
url	http://sfxeu10.hosted.exlibrisgroup.com/loughborough?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-01T18%3A35%3A58IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest_ieee_&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=Beyond%20Bilinear:%20Generalized%20Multimodal%20Factorized%20High-Order%20Pooling%20for%20Visual%20Question%20Answering&rft.jtitle=IEEE%20transaction%20on%20neural%20networks%20and%20learning%20systems&rft.au=Yu,%20Zhou&rft.date=2018-12-01&rft.volume=29&rft.issue=12&rft.spage=5947&rft.epage=5959&rft.pages=5947-5959&rft.issn=2162-237X&rft.eissn=2162-2388&rft.coden=ITNNAL&rft_id=info:doi/10.1109/TNNLS.2018.2817340&rft_dat=%3Cproquest_ieee_%3E2068342533%3C/proquest_ieee_%3E%3Cgrp_id%3Ecdi_FETCH-LOGICAL-c417t-c87921c23edad465ef24b7b9883c1541e509331db8360e291402691ba1dd38413%3C/grp_id%3E%3Coa%3E%3C/oa%3E%3Curl%3E%3C/url%3E&rft_id=info:oai/&rft_pqid=2137588726&rft_id=info:pmid/29993847&rft_ieee_id=8334194&rfr_iscdi=true