Please use this identifier to cite or link to this item:
http://repository.iiitd.edu.in/xmlui/handle/123456789/2078Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Goel, Palaash | - |
| dc.contributor.author | Akhtar, Md. Shad (Advisor) | - |
| dc.date.accessioned | 2026-09-02T10:06:20Z | - |
| dc.date.available | 2026-09-02T10:06:20Z | - |
| dc.date.issued | 2024-11-28 | - |
| dc.identifier.uri | http://repository.iiitd.edu.in/xmlui/handle/123456789/2078 | - |
| dc.description.abstract | Multimodal Sarcasm Explanation (MuSE) is a challenging natural language understanding task that deals with training machines to understand the semantic incongruence present in sarcastic social media posts comprising of an image and a corresponding text caption and generating a natural language explanation to reveal the implicit (hidden) meaning behind them. We use the MORE dataset for the same. The current state of the art for this task (14) makes use of a ‘multi-source semantic graph’ which incor porates external world knowledge along with concepts extracted from both the images and their corre sponding captions to facilitate the reasoning process and lead to a better explanation model. After careful analysing of some of their limitations, we proposed the novel Target-aUgmented shaRed fusion-Based sarcasm explanatiOn model, aka. TURBO, for the task of MuSE that: 1. Utilizes a knowledge graph to incorporate external knowledge, similar to TEAM, while overcom ing the aforementioned limitations 2. Incorporates a novel shared fusion mechanism for learning important information from both the visual and textual modalities 3. Utilizes manually annotated information about the target of sarcasm in our model 4. Beats the current state-of-the-art by roughly 2-3% on average It is important to note that TURBOfixes a problem in the implementation of the previous version of this model (presented as part of last semester’s work). Additionally, we replace the term “cause of sarcasm” with “target of sarcasm” since the latter more appropriately represents the purpose/meaning of the annotation and is thus, easier to understand as well. | en_US |
| dc.language.iso | en_US | en_US |
| dc.publisher | IIIT-Delhi | en_US |
| dc.subject | Multimodal Sarcasm Explanation | en_US |
| dc.subject | Multimodality Analysis | en_US |
| dc.subject | Natural Language Processing | en_US |
| dc.subject | Knowledge Graph | en_US |
| dc.title | Target-augmented shared fusion based multimodal sarcasm explanation generation | en_US |
| dc.type | Other | en_US |
| Appears in Collections: | Year-2024 | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| BTPReport_2021547_PalaashGoel - Palaash Goel.pdf Restricted Access | 1.35 MB | Adobe PDF | View/Open Request a copy |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.