awesome-multimodal-ml
Reading list for research topics in multimodal machine learning
https://github.com/pliang279/awesome-multimodal-ml
Last synced: 5 days ago
JSON representation
-
Applications and Datasets
-
Affect Recognition and Multimodal Language
- End-to-end Facial and Physiological Model for Affective Computing and Applications
- Affective Computing for Large-Scale Heterogeneous Multimedia Data: A Survey
- Towards Multimodal Sarcasm Detection (An Obviously_Perfect Paper)
- Multi-modal Approach for Affective Computing
- Multimodal Language Analysis with Recurrent Multistage Fusion
- Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph - MultimodalSDK)
- End-to-End Multimodal Emotion Recognition using Deep Neural Networks
- Decoding Children’s Social Behavior
- Collecting Large, Richly Annotated Facial-Expression Databases from Movies
- The Interactive Emotional Dyadic Motion Capture (IEMOCAP) Database
- AMHUSE - A Multimodal dataset for HUmor SEnsing
- Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph - MultimodalSDK)
- Collecting Large, Richly Annotated Facial-Expression Databases from Movies
- Multi-attention Recurrent Network for Human Communication Comprehension - MultimodalSDK)
-
Audio and Visual
- Music Gesture for Visual Sound Separation
- Co-Compressing and Unifying Deep CNN Models for Efficient Human Face and Speaker Recognition
- Learning Individual Styles of Conversational Gesture
- Capture, Learning, and Synthesis of 3D Speaking Styles
- Disjoint Mapping Network for Cross-modal Matching of Voices and Faces
- Wav2Pix: Speech-conditioned Face Generation using Generative Adversarial Networks - upc.github.io/wav2pix/)
- Learning Affective Correspondence between Music and Image
- Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input - Discovering-Visual-Objects-and-Spoken-Words)
- Seeing Voices and Hearing Faces: Cross-modal Biometric Matching - nagrani/SVHF-Net)
- Learning to Separate Object Sounds by Watching Unlabeled Video
- Deep Audio-Visual Speech Recognition
- Look, Listen and Learn
- Unsupervised Learning of Spoken Language with Visual Context
- SoundNet: Learning Sound Representations from Unlabeled Video
- Capture, Learning, and Synthesis of 3D Speaking Styles
- Capture, Learning, and Synthesis of 3D Speaking Styles
- Unsupervised Learning of Spoken Language with Visual Context
- Learning to Separate Object Sounds by Watching Unlabeled Video
- Look, Listen and Learn
-
Autonomous Driving
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- nuScenes: A multimodal dataset for autonomous driving
- Multimodal End-to-End Autonomous Driving
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
-
Commonsense Reasoning
- Adventures in Flatland: Perceiving Social Interactions Under Physical Dynamics
- A Logical Model for Supporting Social Commonsense Knowledge Acquisition
- Heterogeneous Graph Learning for Visual Commonsense Reasoning
- SocialIQA: Commonsense Reasoning about Social Interactions
- From Recognition to Cognition: Visual Commonsense Reasoning
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
-
Finance
- A Multimodal Event-driven LSTM Model for Stock Prediction Using Online News
- Multimodal Deep Learning for Finance: Integrating and Forecasting International Stock Markets
- Multimodal deep learning for short-term stock volatility prediction
- Self-Supervised Learning in Event Sequences: A Comparative Study and Hybrid Approach of Generative Modeling and Contrastive Learning
-
Healthcare
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images
- PET-Guided Attention Network for Segmentation of Lung Tumors from PET/CT Images
- Pathomic Fusion: An Integrated Framework for Fusing Histopathology and Genomic Features for Cancer Diagnosis and Prognosis
- Leveraging Medical Visual Question Answering with Supporting Facts
- Unsupervised Multimodal Representation Learning across Medical Images and Reports
- Multimodal Medical Image Retrieval based on Latent Topic Modeling
- Improving Hospital Mortality Prediction with Medical Named Entities and Multimodal Learning
- Knowledge-driven Generative Subspaces for Modeling Multi-view Dependencies in Medical Data
- Multimodal Depression Detection: Fusion Analysis of Paralinguistic, Head Pose and Eye Gaze Behaviors
- Learning the Joint Representation of Heterogeneous Temporal Events for Clinical Endpoint Prediction
- Understanding Coagulopathy using Multi-view Data in the Presence of Sub-Cohorts: A Hierarchical Subspace Approach
- Machine Learning in Multimodal Medical Imaging
- Cross-modal Recurrent Models for Weight Objective Prediction from Multimodal Time-series Data
- SimSensei Kiosk: A Virtual Human Interviewer for Healthcare Decision Support
- Dyadic Behavior Analysis in Depression Severity Assessment Interviews
- Audiovisual Behavior Descriptors for Depression Assessment
-
Human AI Interaction
- Multimodal Human Computer Interaction: A Survey
- Building a multimodal human-robot interface
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Affective multimodal human-computer interaction
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Building a multimodal human-robot interface
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
- Multimodal Human Computer Interaction: A Survey
-
Language and Audio
- Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions
- Lattice Transformer for Speech Translation
- Exploring Phoneme-Level Speech Representations for End-to-End Speech Translation
- Audio Caption: Listen and Tell
- Audio-Linguistic Embeddings for Spoken Sentences
- From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings
- From Audio to Semantics: Approaches To End-to-end Spoken Language Understanding
- Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning
- Deep Voice 2: Multi-Speaker Neural Text-to-Speech
- Deep Voice: Real-time Neural Text-to-Speech
- Text-to-Speech Synthesis
-
Language and Visual QA
- TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation
- Learning to Answer Questions in Dynamic Audio-Visual Scenarios
- SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events - TrafficQA)
- MultiModalQA: complex question answering over text, tables and images
- ManyModalQA: Modality Disambiguation and QA over Diverse Inputs
- Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
- Interactive Language Learning by Question Answering - eric-yuan/qait_public)
- Fusion of Detected Objects in Text for Visual Question Answering
- RUBi: Reducing Unimodal Biases in Visual Question Answering
- GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
- OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
- MUREL: Multimodal Relational Reasoning for Visual Question Answering
- Social-IQ: A Question Answering Benchmark for Artificial Social Intelligence - IQ)
- Probabilistic Neural-symbolic Models for Interpretable Visual Question Answering - clevr)
- Learning to Count Objects in Natural Images for Visual Question Answering - counting)
- Overcoming Language Priors in Visual Question Answering with Adversarial Regularization
- Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding - vqa)
- RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes
- TVQA: Localized, Compositional Video Question Answering
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
- Stacked Latent Attention for Multimodal Reasoning
- Learning to Reason: End-to-End Module Networks for Visual Question Answering
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning - iep) [[dataset generation]](https://github.com/facebookresearch/clevr-dataset-gen)
- Are You Smarter Than A Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding - mcb)
- MovieQA: Understanding Stories in Movies through Question-Answering
- VQA: Visual Question Answering
- Stacked Latent Attention for Multimodal Reasoning
- Are You Smarter Than A Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension
- TVQA: Localized, Compositional Video Question Answering
-
Language Grouding in Navigation
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation - RIPL/robo-vln), [[video]](https://www.youtube.com/watch?v=y16x9n_zP_4), [[project page]](https://zubair-irshad.github.io/projects/robo-vln.html)
- Improving Vision-and-Language Navigation with Image-Text Pairs from the Web
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Vision-and-Dialog Navigation
- Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
- Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
- Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments - lab/touchdown)
- Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
- The Regretful Navigation Agent for Vision-and-Language Navigation - agent)
- Tactical Rewind: Self-Correction via Backtracking in Vision-and-Language Navigation
- Multi-modal Discriminative Model for Vision-and-Language Navigation - RoboNLP Workshop 2019
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation - agent)
- From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following
- Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
- Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout - EnvDrop)
- Attention Based Natural Language Grounding by Navigating Virtual Environment
- Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction - lab/ciff)
- Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
- Embodied Question Answering
- Look Before You Leap: Bridging Model-Free and Model-Based Reinforcement Learning for Planned-Ahead Vision-and-Language Navigation
- Multi-modal Discriminative Model for Vision-and-Language Navigation - RoboNLP Workshop 2019
- Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout - EnvDrop)
-
Language Grounding in Vision
- Core Challenges in Embodied Vision-Language Planning
- Grounding 'Grounding' in NLP
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes - memes-challenge-and-data-set/)
- What Does BERT with Vision Look At?
- Visual Grounding in Video for Unsupervised Word Translation - grounding)
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Grounded Video Description
- Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval - to-Clip_Retrieval)
- Binary Image Selection (BISON): Interpretable Evaluation of Visual Grounding - image-selection)
- Finding “It”: Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos
- SCAN: Learning Hierarchical Compositional Visual Concepts
-
Programming Languages
Categories
Sub Categories
Multimodal Content Generation
112
Human AI Interaction
59
Multimodal Representations
41
Language and Visual QA
31
Multimodal Machine Translation
29
Multimodal Fusion
26
Language Grounding in Vision
26
Language Grouding in Navigation
24
Media Description
21
Robotics
19
Audio and Visual
19
Multi-agent Communication
16
Healthcare
16
Affect Recognition and Multimodal Language
14
Multimodal Alignment
11
Language and Audio
11
Multimodal Reinforcement Learning
11
Generative Learning
10
Bias and Fairness
10
Multimodal Pretraining
10
Multimodal Dialog
8
Knowledge Graphs and Knowledge Bases
7
Analysis of Multimodal Models
6
Multimodal Memory
6
Self-supervised Learning
6
Multimodal Translation
6
Commonsense Reasoning
6
Missing or Imperfect Modalities
5
Few-Shot Learning
5
Language Models
5
Crossmodal Retrieval
5
Multimodal Co-learning
4
Autonomous Driving
4
Human in the Loop Learning
4
Finance
4
Semi-supervised Learning
4
Intepretable Learning
4
Multimodal Transformers
4
Visual, IMU and Wireless
3
Video Generation from Text
3
Adversarial Attacks
3