{"id":18887124,"url":"https://github.com/thunlp-mt/streamingbench","last_synced_at":"2026-03-15T22:13:16.710Z","repository":{"id":261588208,"uuid":"883722360","full_name":"THUNLP-MT/StreamingBench","owner":"THUNLP-MT","description":"StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding","archived":false,"fork":false,"pushed_at":"2025-03-27T03:10:28.000Z","size":32892,"stargazers_count":114,"open_issues_count":3,"forks_count":3,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-03-31T08:12:02.931Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/THUNLP-MT.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-11-05T13:12:54.000Z","updated_at":"2025-03-30T08:46:33.000Z","dependencies_parsed_at":"2025-03-31T08:21:05.172Z","dependency_job_id":null,"html_url":"https://github.com/THUNLP-MT/StreamingBench","commit_stats":null,"previous_names":["thunlp-mt/streamingbench"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FStreamingBench","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FStreamingBench/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FStreamingBench/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FStreamingBench/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/THUNLP-MT","download_url":"https://codeload.github.com/THUNLP-MT/StreamingBench/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247622988,"owners_count":20968575,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-08T07:34:19.780Z","updated_at":"2026-03-15T22:13:16.703Z","avatar_url":"https://github.com/THUNLP-MT.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/icon.png\" width=\"100%\" alt=\"StreamingBench Banner\"\u003e\n\n  \u003cdiv style=\"margin: 30px 0\"\u003e\n    \u003ca href=\"https://streamingbench.github.io/\" style=\"margin: 0 10px\"\u003e🏠 Project Page\u003c/a\u003e |\n    \u003ca href=\"https://arxiv.org/abs/2411.03628\" style=\"margin: 0 10px\"\u003e📄 arXiv Paper\u003c/a\u003e |\n    \u003ca href=\"https://huggingface.co/datasets/mjuicem/StreamingBench\" style=\"margin: 0 10px\"\u003e📦 Dataset\u003c/a\u003e |\n    \u003ca href=\"https://streamingbench.github.io/#leaderboard\" style=\"margin: 0 10px\"\u003e🏅Leaderboard\u003c/a\u003e\n  \u003c/div\u003e\n\u003c/div\u003e\n\n**StreamingBench** evaluates **Multimodal Large Language Models (MLLMs)** in real-time, streaming video understanding tasks. 🌟\n\n\n\n\n------\n\n[**NEW!** 2025.05.15] 🔥: [Seed1.5-VL](https://github.com/ByteDance-Seed/Seed1.5-VL) achieved ALL model SOTA with a score of 82.80 on the Proactive Output.\n\n[**NEW!** 2025.03.17] ⭐: [ViSpeeker](https://arxiv.org/abs/2503.12769) achieved Open-Source SOTA with a score of 61.60 on the Omni-Source Understanding.\n\n[**NEW!** 2025.01.14] 🚀: [MiniCPM-o 2.6](https://github.com/OpenBMB/MiniCPM-o) achieved Streaming SOTA with a score of 66.01 on the Overall benchmark.\n\n[**NEW!** 2025.01.06] 🏆: [Dispider](https://github.com/Mark12Ding/Dispider) achieved Streaming SOTA with a score of 53.12 on the Overall benchmark.\n\n[**NEW!** 2024.12.09] 🎉: [InternLM-XComposer2.5-OmniLive](https://github.com/InternLM/InternLM-XComposer) achieved 73.79 on Real-Time Visual Understanding.\n\n------\n\n## 🎞️ Overview\n\nAs MLLMs continue to advance, they remain largely focused on offline video comprehension, where all frames are pre-loaded before making queries. However, this is far from the human ability to process and respond to video streams in real-time, capturing the dynamic nature of multimedia content. To bridge this gap, **StreamingBench** introduces the first comprehensive benchmark for streaming video understanding in MLLMs.\n\n### Key Evaluation Aspects\n- 🎯 **Real-time Visual Understanding**: Can the model process and respond to visual changes in real-time?\n- 🔊 **Omni-source Understanding**: Does the model integrate visual and audio inputs synchronously in real-time video streams?\n- 🎬 **Contextual Understanding**: Can the model comprehend the broader context within video streams?\n\n### Dataset Statistics\n- 📊 **900** diverse videos\n- 📝 **4,500** human-annotated QA pairs\n- ⏱️ Five questions per video at different timestamps\n#### 🎬 Video Categories\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/StreamingBench_Video.png\" width=\"80%\" alt=\"Video Categories\"\u003e\n\u003c/div\u003e\n\n#### 🔍 Task Taxonomy\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/task_taxonomy.png\" width=\"80%\" alt=\"Task Taxonomy\"\u003e\n\u003c/div\u003e\n\n## 📐 Dataset Examples\nhttps://github.com/user-attachments/assets/e6d1655d-ab3f-47a7-973a-8fd6c8962307\n\u003cdiv align=\"center\"\u003e\n  \u003cvideo width=\"100%\" controls\u003e\n    \u003csource src=\"./figs/example.video\" type=\"video/mp4\"\u003e\n    Your browser does not support the video tag.\n  \u003c/video\u003e\n\u003c/div\u003e\n\n## 🔮 Evaluation Pipeline\n\n### Requirements\n\n- Python 3.x\n- ffmpeg-python\n\n### Data Preparation\n\n1. **Download Dataset**: Retrieve all necessary files from the [StreamingBench Dataset](https://huggingface.co/datasets/mjuicem/StreamingBench).\n   \n2. **Decompress Files**: Extract the downloaded files and organize them in the `./data` directory as follows:\n\n   ```\n   StreamingBench/\n   ├── data/\n   │   ├── real/               # Unzip Real Time Visual Understanding_*.zip into this folder\n   │   ├── omni/               # Unzip other .zip files into this folder\n   │   ├── sqa/                # Unzip Sequential Question Answering_*.zip into this folder\n   │   └── proactive/          # Unzip Proactive Output_*.zip into this folder\n   ```\n\n3. **Preprocess Data**: Run the following command to preprocess the data:\n\n   ```bash\n   cd ./scripts\n   bash preprocess.sh\n   ```\n\n### Model Preparation\n\nPrepare your own model for evaluation by following the instructions provided [here](./docs/model_guide.md). This guide will help you set up and configure your model to ensure it is ready for testing against the dataset.\n\n### Evaluation\n\nNow you can run the benchmark:\n\n```sh\nbash eval.sh\n```\n\nThis will run the benchmark and save the results to the specified output file. Then you can calculate the metrics using the following command:\n```sh\nbash stats.sh\n```\n\n## 🔬 Experimental Results\n\n### Performance of Various MLLMs on StreamingBench\n- 60 seconds of context preceding the query time (Main)\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/result_2.png\" width=\"80%\" alt=\"Task Taxonomy\"\u003e\n\u003c/div\u003e\n\n- All Context (+ Long Context)\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/result_1.png\" width=\"80%\" alt=\"Task Taxonomy\"\u003e\n\u003c/div\u003e\n\n\n\n- Comparison of Main Experiment vs. 60 Seconds of Video Context\n- \u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/heatmap.png\" width=\"80%\" alt=\"Task Taxonomy\"\u003e\n\u003c/div\u003e\n\n### Performance of Different MLLMs on the Proactive Output Task\n*\"≤ xs\" means that the answer is considered correct if the actual output time is within x seconds of the ground truth.*\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"./figs/po.png\" width=\"80%\" alt=\"Task Taxonomy\"\u003e\n\u003c/div\u003e\n\n\n## 📝 Citation\n```bibtex\n@article{lin2024streaming,\n  title={StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding},\n  author={Junming Lin and Zheng Fang and Chi Chen and Zihao Wan and Fuwen Luo and Peng Li and Yang Liu and Maosong Sun},\n  journal={arXiv preprint arXiv:2411.03628},\n  year={2024}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fthunlp-mt%2Fstreamingbench","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fthunlp-mt%2Fstreamingbench","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fthunlp-mt%2Fstreamingbench/lists"}