https://github.com/tiger-ai-lab/viescore
Visual Instruction-guided Explainable Metric. Code for "Towards Explainable Metrics for Conditional Image Synthesis Evaluation" (ACL 2024 main)
https://github.com/tiger-ai-lab/viescore
computer-vision gpt4vision image-editing image-generation visual-question-answering
Last synced: about 1 year ago
JSON representation
Visual Instruction-guided Explainable Metric. Code for "Towards Explainable Metrics for Conditional Image Synthesis Evaluation" (ACL 2024 main)
- Host: GitHub
- URL: https://github.com/tiger-ai-lab/viescore
- Owner: TIGER-AI-Lab
- License: mit
- Created: 2023-12-22T02:20:44.000Z (over 2 years ago)
- Default Branch: main
- Last Pushed: 2024-11-19T01:46:51.000Z (over 1 year ago)
- Last Synced: 2024-11-19T02:31:33.996Z (over 1 year ago)
- Topics: computer-vision, gpt4vision, image-editing, image-generation, visual-question-answering
- Language: Python
- Homepage: https://tiger-ai-lab.github.io/VIEScore/
- Size: 21.6 MB
- Stars: 28
- Watchers: 4
- Forks: 1
- Open Issues: 0
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
README
# VIEScore
[](https://arxiv.org/abs/2312.14867)
[](https://github.com/TIGER-AI-Lab/VIEScore/graphs/contributors)
[](https://github.com/TIGER-AI-Lab/VIEScore/blob/main/LICENSE)
[](https://github.com/TIGER-AI-Lab/VIEScore)
[](https://hits.seeyoufarm.com)
This repository hosts the code and data of our ACL 2024 Paper [VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation](https://tiger-ai-lab.github.io/VIEScore/).
VIEScore is a Visual Instruction-guided Explainable metric for evaluating any conditional image generation tasks.
Metrics in the future would provide the score and the rationale, enabling the understanding of each judgment. Which method (VIEScore or traditional metrics) is βcloserβ to the human perspective?
## π° News
* 2024 Jun 17: We released the standalone version of VIEScore.
* 2024 May 23: We released all the results and notebook to visualize the results.
* 2024 May 23: Added Gemini-1.5-pro results.
* 2024 May 16: Added GPT4o results and we found that GPT4o achieve on par correlation with human across all tasks!
* 2024 May 15: VIEScore is accepted to ACL2024 (main)!
* 2024 Jan 11: Code is released!
* 2023 Dec 24: Paper available on [Arxiv](https://arxiv.org/abs/2312.14867). Code coming Soon!

> VIEScore gives an SC(semantic consistency score), PQ(perceptual quality score), and O (Overall score) to evaluate your image/video.
## Paper implementation
See https://github.com/TIGER-AI-Lab/VIEScore/tree/main/paper_implementation
```python
$ python3 run.py --help
usage: run.py [-h] [--task {tie,mie,t2i,cig,sdig,msdig,sdie}] [--mllm {gpt4v, gpt4o, llava,blip2,fuyu,qwenvl,cogvlm,instructblip,openflamingo, gemini}] [--setting {0shot,1shot}] [--context_file CONTEXT_FILE]
[--guess_if_cannot_parse]
Run different task on VIEScore.
optional arguments:
-h, --help show this help message and exit
--task {tie,mie,t2i,cig,sdig,msdig,sdie}
Select the task to run
--mllm {gpt4v, gpt4o, llava,blip2,fuyu,qwenvl,cogvlm,instructblip,openflamingo, gemini}
Select the MLLM model to use
--setting {0shot,1shot}
Select the incontext learning setting
--context_file CONTEXT_FILE
Which context file to use.
--guess_if_cannot_parse
Guess a value if the output cannot be parsed.
```
## Standard Version (For Development and Extension)
See https://github.com/TIGER-AI-Lab/VIEScore/tree/main/viescore
```python
from viescore import VIEScore
backbone = "gemini"
vie_score = VIEScore(backbone=backbone, task="t2v")
score_list = vie_score.evaluate(pil_image, text_prompt)
sementics_score, quality_score, overall_score = score_list
```
## Paper Results
## Citation
Please kindly cite our paper if you use our code, data, models or results:
```bibtex
@misc{ku2023viescore,
title={VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation},
author={Max Ku and Dongfu Jiang and Cong Wei and Xiang Yue and Wenhu Chen},
year={2023},
eprint={2312.14867},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
```