https://github.com/shehanmunasinghe/ai701-project-g02
https://github.com/shehanmunasinghe/ai701-project-g02
Last synced: over 1 year ago
JSON representation
- Host: GitHub
- URL: https://github.com/shehanmunasinghe/ai701-project-g02
- Owner: shehanmunasinghe
- Created: 2023-11-24T11:01:15.000Z (over 2 years ago)
- Default Branch: main
- Last Pushed: 2023-11-25T15:55:12.000Z (over 2 years ago)
- Last Synced: 2025-03-15T10:11:23.474Z (over 1 year ago)
- Language: Python
- Size: 104 KB
- Stars: 0
- Watchers: 1
- Forks: 0
- Open Issues: 0
-
Metadata Files:
- Readme: README.md
- Support: support_long_term_videos/adjacent_frame_merge2.py
Awesome Lists containing this project
README
# Extending Video-based Large Multimodal Models
Shehan Munasinghe, Rusiru Thushara, Mohamed Insaf Ismithdeen
Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
*{shehan.munasinghe, rusiru.achchige, mohamed.ismithdeen}@mbzuai.ac.ae*
AI701 Project : Group ID = G-02
---
## 1 - Setup and Demo
See [here](1-CLI_DEMO.md) for instrunctions on insallation and running the CLI demo.
## 2 - Training
See [here](2-Training.md) for more details.
## 3 - Quantitative Evaluation Framework for Video-based Conversational Models
See [here](3-Benchmark.md) for more details.
## 4 - Quantitative Evaluation of Conversation-based Video Spatial Grounding
See [here](4-GroundingBenchmark.md) for more details.