107 - Multi-Modal Transformers, With Hao Tan And Mohit Bansal NLP Highlights podcast

Artwork

Artificial Intelligence Tech Science NLP Highlights Allen Institute for Artificial Intelligence Tell Us

Indhold leveret af NLP Highlights and Allen Institute for Artificial Intelligence. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af NLP Highlights and Allen Institute for Artificial Intelligence eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.

NLP Highlights « »
107 - Multi-Modal Transformers, with Hao Tan and Mohit Bansal

5y ago 37:34

Del

MP3•Episode hjem

Indhold leveret af NLP Highlights and Allen Institute for Artificial Intelligence. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af NLP Highlights and Allen Institute for Artificial Intelligence eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.

In this episode, we invite Hao Tan and Mohit Bansal to talk about multi-modal training of transformers, focusing in particular on their EMNLP 2019 paper that introduced LXMERT, a vision+language transformer. We spend the first third of the episode talking about why you might want to have multi-modal representations. We then move to the specifics of LXMERT, including the model structure, the losses that are used to encourage cross-modal representations, and the data that is used. Along the way, we mention latent alignments between images and captions, the granularity of captions, and machine translation even comes up a few times. We conclude with some speculation on the future of multi-modal representations. Hao's website: http://www.cs.unc.edu/~airsplay/ Mohit's website: http://www.cs.unc.edu/~mbansal/ LXMERT paper: https://www.aclweb.org/anthology/D19-1514/

… continue reading

145 episoder

#Artificial Intelligence #Tech #Science #NLP Highlights #Allen Institute for Artificial Intelligence #Tell Us

Artwork

107 - Multi-Modal Transformers, with Hao Tan and Mohit Bansal

286 subscribers

published 5y ago

Del

MP3•Episode hjem

Indhold leveret af NLP Highlights and Allen Institute for Artificial Intelligence. Alt podcastindhold inklusive episoder, grafik og podcastbeskrivelser uploades og leveres direkte af NLP Highlights and Allen Institute for Artificial Intelligence eller deres podcastplatformspartner. Hvis du mener, at nogen bruger dit ophavsretligt beskyttede værk uden din tilladelse, kan du følge processen beskrevet her https://da.player.fm/legal.

In this episode, we invite Hao Tan and Mohit Bansal to talk about multi-modal training of transformers, focusing in particular on their EMNLP 2019 paper that introduced LXMERT, a vision+language transformer. We spend the first third of the episode talking about why you might want to have multi-modal representations. We then move to the specifics of LXMERT, including the model structure, the losses that are used to encourage cross-modal representations, and the data that is used. Along the way, we mention latent alignments between images and captions, the granularity of captions, and machine translation even comes up a few times. We conclude with some speculation on the future of multi-modal representations. Hao's website: http://www.cs.unc.edu/~airsplay/ Mohit's website: http://www.cs.unc.edu/~mbansal/ LXMERT paper: https://www.aclweb.org/anthology/D19-1514/

… continue reading

145 episoder

#Artificial Intelligence #Tech #Science #NLP Highlights #Allen Institute for Artificial Intelligence #Tell Us

Alle episoder

×

Velkommen til Player FM!

Player FM is scanning the web for high-quality podcasts for you to enjoy right now. It's the best podcast app and works on Android, iPhone, and the web. Signup to sync subscriptions across devices.

Lyt til 500+ emner

Lyt til dette show, mens du udforsker