{"id":19763878,"url":"https://github.com/khuangaf/chocolate","last_synced_at":"2026-03-01T23:31:19.986Z","repository":{"id":212844609,"uuid":"732088170","full_name":"khuangaf/CHOCOLATE","owner":"khuangaf","description":"Code and data for the ACL 2024 Findings paper \"Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning\"","archived":false,"fork":false,"pushed_at":"2024-06-05T04:37:42.000Z","size":3653,"stargazers_count":26,"open_issues_count":1,"forks_count":1,"subscribers_count":3,"default_branch":"master","last_synced_at":"2025-04-30T14:33:22.422Z","etag":null,"topics":["chart-captioning","chart-summarization","chart-understanding","factuality","faithfulness","large-vision-language-models"],"latest_commit_sha":null,"homepage":"https://khuangaf.github.io/CHOCOLATE","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/khuangaf.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-12-15T16:02:19.000Z","updated_at":"2025-04-22T12:35:56.000Z","dependencies_parsed_at":"2024-02-11T20:01:34.737Z","dependency_job_id":"5a0e7066-c18e-47ec-b420-d6a938af062d","html_url":"https://github.com/khuangaf/CHOCOLATE","commit_stats":null,"previous_names":["khuangaf/chocolate"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/khuangaf/CHOCOLATE","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/khuangaf%2FCHOCOLATE","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/khuangaf%2FCHOCOLATE/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/khuangaf%2FCHOCOLATE/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/khuangaf%2FCHOCOLATE/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/khuangaf","download_url":"https://codeload.github.com/khuangaf/CHOCOLATE/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/khuangaf%2FCHOCOLATE/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29987698,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-01T22:42:38.399Z","status":"ssl_error","status_checked_at":"2026-03-01T22:41:51.863Z","response_time":124,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["chart-captioning","chart-summarization","chart-understanding","factuality","faithfulness","large-vision-language-models"],"created_at":"2024-11-12T04:11:26.508Z","updated_at":"2026-03-01T23:31:19.961Z","avatar_url":"https://github.com/khuangaf.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# [ACL 2024 Findings] Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning\n\n\u003cdiv align=\"center\"\u003e\n\u003ca href=\"https://khuangaf.github.io/\"\u003eKung-Hsiang Huang\u003c/a\u003e†, Mingyang Zhou*, Hou Pong Chan‡,\nYi R. Fung†, Zhenhailong Wang†, Lingyu Zhang*, Shih-Fu Chang*, Heng Ji†\n\n\u003c/div\u003e\n\u003cdiv align=\"center\"\u003e\n\u003cstrong\u003eUniversity of Illinois Urbana-Champaign†\u003c/strong\u003e\n\n\u003cstrong\u003eColumbia University*\u003c/strong\u003e\n\u003cstrong\u003eUniversity of Macau‡\u003c/strong\u003e\n\u003c/div\u003e\n\n\u003cdiv align=\"center\"\u003e\n\u003chr\u003e\n\n\u003c!-- [![arXiv](https://img.shields.io/badge/arXiv-2312.10160-b31b1b.svg?style=for-the-badge)](https://arxiv.org/abs/2312.10160) --\u003e\n\n\u003ca href='https://arxiv.org/abs/2312.10160'\u003e\u003cimg src='https://img.shields.io/badge/arXiv-2312.10160-b31b1b.svg'\u003e\u003c/a\u003e \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\n\u003ca href='https://khuangaf.github.io/CHOCOLATE/'\u003e\u003cimg src='https://img.shields.io/badge/Project-Page-Green'\u003e\u003c/a\u003e \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\n\u003ca href='https://github.com/khuangaf/CHOCOLATE?tab=Apache-2.0-1-ov-file#readme'\u003e\u003cimg src='https://img.shields.io/badge/License-Apache_2.0-blue'\u003e\u003c/a\u003e \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\n\u003ca href='https://paperswithcode.com/dataset/chocolate'\u003e\u003cimg src='https://img.shields.io/badge/Paper_With_Code-CHOCOLATE-48e5e8'\u003e\u003c/a\u003e \n\n[![ChartVE](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-ChartVE-blue)](https://huggingface.co/khhuang/chartve) \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; \n[![Chart-to-Table](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Chart_to_Table-blue)](https://huggingface.co/khhuang/chart-to-table) \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; \n[![CHOCOLATE](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-CHOCOLATE-blue)](https://huggingface.co/datasets/khhuang/CHOCOLATE)\n\u003c/div\u003e\n\n\nThis repository holds the CHOCOLATE benchmark for assessing the factuality of chart captioning systems and facilitating the Chart Caption Factual Error Correction task. The dataset includes an error analysis for six different models on two distinct datasets. These includes:\n\n* LVLM: GPT-4V, Bard (before Gemini)\n* LLM-based Pipeline: DePlot + GPT-4\n* Fine-tuned Model: ChartT5, MatCha, UniChart\n\nAnnotations are conducted on the VisText and Chart-to-Text (pew split) datasets. This ensures a wide range of data and types of factual errors. For more information, please visit our [project page](https://khuangaf.github.io/CHOCOLATE).\n\nResults are shown in the below figure and table. We found that all captioning models often generate captions that are factually inconsistent with the input chart. In fact, even for highly capable **LVLMs, their non-factual rate is a whopping 81.27%**.\n\n\u003cimg src=\"./error_distribution.png\"  class=\"center\"\u003e\n\n\u003cdiv align=\"center\"\u003e\n\u003ctable\u003e\n  \u003ctr\u003e\n    \u003cth style=\"font-weight: bold;\"\u003e\u003c/th\u003e\n    \u003cth style=\"font-weight: bold;\" colspan=\"2\"\u003eCHOCOLATE-LVLM\u003c/th\u003e\n    \u003cth style=\"font-weight: bold;\" colspan=\"2\"\u003eCHOCOLATE-LLM\u003c/th\u003e\n    \u003cth style=\"font-weight: bold;\" colspan=\"2\"\u003eCHOCOLATE-FT\u003c/th\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003c/td\u003e\n    \u003ctd\u003e# Factual\u003c/td\u003e\n    \u003ctd\u003e# Non-factual\u003c/td\u003e\n    \u003ctd\u003e# Factual\u003c/td\u003e\n    \u003ctd\u003e# Non-factual\u003c/td\u003e\n    \u003ctd\u003e# Factual\u003c/td\u003e\n    \u003ctd\u003e# Non-factual\u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003eSentence\u003c/td\u003e\n    \u003ctd\u003e1,683\u003c/td\u003e\n    \u003ctd\u003e1,270\u003c/td\u003e\n    \u003ctd\u003e518\u003c/td\u003e\n    \u003ctd\u003e469\u003c/td\u003e\n    \u003ctd\u003e360\u003c/td\u003e\n    \u003ctd\u003e1,023\u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003eCaption\u003c/td\u003e\n    \u003ctd\u003e74\u003c/td\u003e\n    \u003ctd\u003e321\u003c/td\u003e\n    \u003ctd\u003e27\u003c/td\u003e\n    \u003ctd\u003e169\u003c/td\u003e\n    \u003ctd\u003e112\u003c/td\u003e\n    \u003ctd\u003e484\u003c/td\u003e\n  \u003c/tr\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\n\n## Spotlights\n\n* CHOCOLATE - The first factuality benchmark for chart captioning.\n* CHOCOLATE is also used to establish the Chart Caption Factual Error Correction task.\n* Comming soon\n    - [x] The CHOCOLATE benchmark\n    - [x] The ChartVE metric ([khhuang/chartve](https://huggingface.co/khhuang/chartve))\n    - [x] The Chart-To-Table model ([khhuang/chart-to-table](https://huggingface.co/khhuang/chart-to-table))\n    - [ ] Scripts for table-based error correction \n    - [ ] Evaluation scripts\n          \n\n## The CHOCOLATE Benchmark\n\nWe release the data for the CHOCOLATE benchmark at `data/chocolate.json`. CHOCOLATE is also available on [HuggingFace](https://huggingface.co/datasets/khhuang/CHOCOLATE)🤗. \n\n### Data Structure\n\nEach instance in the json file corresponds to an annotation for a generated caption. Below, we illustrate the fields within each instance:\n\n* **sentences**: A list of caption sentences.\n* **labels**: A list of list, where the outer list correspond to sentence and the inner list corresponds to the errors within each sentence.\n* **model**: A string that represents the model producing the caption.\n* **dataset**: A string that represents which dataset the chart was sampled from.\n* **image_path**: An URL to the chart image.\n* **_id**: A unique identifier for this instance.\n\n## ChartVE\n\n\nChartVE is a visual entailment model for evaluating the factuality of a generated caption sentence with regard to the input chart. The model takes in a chart figure and a caption sentence as input, and outputs an entailment probability. The underlying architecture of this model is UniChart.\n\nNote that this model expects a caption sentence as textual inputs. For captions that are longer than one sentences, one should split the caption into multiple sentences, feed individual sentences to ChartVE, and then aggregate the scores. Below, we provide an example of how to use ChartVE.\n\n```python\nfrom transformers import DonutProcessor, VisionEncoderDecoderModel\nfrom PIL import Image\n\nmodel_name = \"khhuang/chartve\"\nmodel = VisionEncoderDecoderModel.from_pretrained(model_name).cuda()\nprocessor = DonutProcessor.from_pretrained(model_name)\n\nimage_path = \"PATH_TO_IMAGE\"\n\ndef format_query(sentence):\n    return f\"Does the image entails this statement: \\\"{sentence}\\\"?\"\n\n# Format text inputs\nCAPTION_SENTENCE = \"The state that has the highest number of population is California.\"\nquery = format_query(CAPTION_SENTENCE)\n\n# Encode chart figure and tokenize text\nimg = Image.open(IMAGE_PATH)\npixel_values = processor(img.convert(\"RGB\"), random_padding=False, return_tensors=\"pt\").pixel_values\npixel_values = pixel_values.cuda()\ndecoder_input_ids = processor.tokenizer(query, add_special_tokens=False, return_tensors=\"pt\", max_length=510).input_ids.cuda()\n\n\noutputs = model(pixel_values, decoder_input_ids=decoder_input_ids)\n\n# positive_logit = outputs['logits'].squeeze()[-1,49922]\n# negative_logit = outputs['logits'].squeeze()[-1,2334] \n\n# Probe the probability of generating \"yes\"\nbinary_entail_prob_positive = torch.nn.functional.softmax(outputs['logits'].squeeze()[-1,[2334, 49922]])[1].item()\n\n# binary_entail_prob_positive corresponds to the computed probability that the chart entails the caption sentence.\n\n```\n\nThe meta-evaluation scripts can be found in [ChartVE Meta-evaluation.ipynb](https://github.com/khuangaf/CHOCOLATE/blob/master/ChartVE%20Meta-evaluation.ipynb).\n\n## C2TFEC\n\nThe proposed C2TFEC framework consists of two components: chart-to-table conversion and table-based error rectification.\n\n### Chart-To-Table\n\nThe Chart-To-Table model ([khhuang/chart-to-table](https://huggingface.co/khhuang/chart-to-table)) is trained to convert a chart into a structured table. The generated tables use \u0026\u0026\u0026 to delimit rows and | to delimit columns. The underlying architecture of this model is UniChart. Below, we provide an example of how to use our Chart-To-Table model.\n\n\n```python\nfrom transformers import DonutProcessor, VisionEncoderDecoderModel\nfrom PIL import Image\n\nmodel_name = \"khhuang/chart-to-table\"\nmodel = VisionEncoderDecoderModel.from_pretrained(model_name).cuda()\nprocessor = DonutProcessor.from_pretrained(model_name)\n\nimage_path = \"PATH_TO_IMAGE\"\n\ndef format_query(sentence):\n    return f\"Does the image entails this statement: \\\"{sentence}\\\"?\"\n\n# Format text inputs\n\ninput_prompt = \"\u003cdata_table_generation\u003e \u003cs_answer\u003e\"\n\n# Encode chart figure and tokenize text\nimg = Image.open(IMAGE_PATH)\npixel_values = processor(img.convert(\"RGB\"), random_padding=False, return_tensors=\"pt\").pixel_values\npixel_values = pixel_values.cuda()\ndecoder_input_ids = processor.tokenizer(input_prompt, add_special_tokens=False, return_tensors=\"pt\", max_length=510).input_ids.cuda()\n\n# Generate a table\noutputs = model.generate(\n        pixel_values.to(device),\n        decoder_input_ids=decoder_input_ids.to(device),\n        max_length=model.decoder.config.max_position_embeddings,\n        early_stopping=True,\n        pad_token_id=processor.tokenizer.pad_token_id,\n        eos_token_id=processor.tokenizer.eos_token_id,\n        use_cache=True,\n        num_beams=4,\n        bad_words_ids=[[processor.tokenizer.unk_token_id]],\n        return_dict_in_generate=True,\n    )\n    \nsequence = processor.batch_decode(outputs.sequences)[0]\nsequence = sequence.replace(processor.tokenizer.eos_token, \"\").replace(processor.tokenizer.pad_token, \"\")\n\n# Extract the data table\nextracted_table = sequence.split(\"\u003cs_answer\u003e\")[1].strip()\n```\n\n\n\n## Citation\n```bibtex\n@inproceedings{huang-etal-2024-lvlms,\n    title = \"Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning\",\n    author = \"Huang, Kung-Hsiang  and\n      Zhou, Mingyang and\n      Chan, Hou Pong  and\n      Fung, Yi R. and\n      Wang, Zhenhailong and\n      Zhang, Lingyu and\n      Chang, Shih-Fu and\n      Ji, Heng\",\n    booktitle = \"Findings of the Association for Computational Linguistics: ACL 2024\",\n    month = aug,\n    year = \"2024\",\n    publisher = \"Association for Computational Linguistics\",\n    url = \"https://aclanthology.org/2023.findings-acl.85\",\n    doi = \"10.18653/v1/2023.findings-acl.85\",\n    pages = \"1314--1326\",\n}    \n```\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkhuangaf%2Fchocolate","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkhuangaf%2Fchocolate","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkhuangaf%2Fchocolate/lists"}