{"id":23636426,"url":"https://github.com/x-izhang/libra","last_synced_at":"2025-06-29T14:31:58.540Z","repository":{"id":265264213,"uuid":"895628273","full_name":"X-iZhang/Libra","owner":"X-iZhang","description":"[ACL 2025] ⚖️ Temporally-aware MLLM for Biomedical Radiology Analysis and Report Generation. Flexible toolkit with LLM backbone support, real-time validation, training resumption, and smart model saving.","archived":false,"fork":false,"pushed_at":"2025-06-15T19:24:04.000Z","size":14288,"stargazers_count":14,"open_issues_count":0,"forks_count":1,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-06-15T21:03:33.546Z","etag":null,"topics":["ai4science","chest-xrays","llama3","medical-image-analysis","multimodal-large-language-models","radiology-report-generation","vision-language-model"],"latest_commit_sha":null,"homepage":"https://x-izhang.github.io/Libra_v1.0/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/X-iZhang.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-11-28T15:07:07.000Z","updated_at":"2025-06-15T19:24:07.000Z","dependencies_parsed_at":"2024-12-19T21:30:34.224Z","dependency_job_id":"f475a90c-dbbe-4e42-a196-b477a39a8954","html_url":"https://github.com/X-iZhang/Libra","commit_stats":null,"previous_names":["x-izhang/libra"],"tags_count":2,"template":false,"template_full_name":null,"purl":"pkg:github/X-iZhang/Libra","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/X-iZhang%2FLibra","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/X-iZhang%2FLibra/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/X-iZhang%2FLibra/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/X-iZhang%2FLibra/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/X-iZhang","download_url":"https://codeload.github.com/X-iZhang/Libra/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/X-iZhang%2FLibra/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":262608870,"owners_count":23336583,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai4science","chest-xrays","llama3","medical-image-analysis","multimodal-large-language-models","radiology-report-generation","vision-language-model"],"created_at":"2024-12-28T06:12:18.580Z","updated_at":"2025-06-29T14:31:58.520Z","avatar_url":"https://github.com/X-iZhang.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003ch1 align=\"center\"\u003e\n  \u003cimg src=\"assets/Libra_logo_c.png\" alt=\"Libra Logo\" width=\"27\" style=\"position: relative; top: -2px;\"/\u003e\n  \u003cstrong\u003eLibra: Leveraging Temporal Images for Biomedical Radiology Analysis\u003c/strong\u003e\n\u003c/h1\u003e\n\n\u003cdiv align=\"center\"\u003e\n\n[![ReXrank](https://img.shields.io/badge/🏆_Libra-Top_Model_on_ReXrank-firebrick)](https://rexrank.ai/)\n[![Project Page](https://img.shields.io/badge/Project-Page-Green?logo=webauthn)](https://x-izhang.github.io/Libra_v1.0/)\n[![Docs](https://img.shields.io/badge/-deepwiki-0A66C2?logo=readthedocs\u0026logoColor=white\u0026color=7289DA\u0026labelColor=grey)](https://deepwiki.com/X-iZhang/Libra)\n[![Demo](https://img.shields.io/badge/⚡-Online%20Demo-yellow.svg)](https://huggingface.co/spaces/X-iZhang/Libra)\n[![hf_space](https://img.shields.io/badge/%F0%9F%A4%97%20-Hugging%20Face-blue)](https://huggingface.co/collections/X-iZhang/libra-6772bfccc6079298a0fa5f8d)\n[![arXiv](https://img.shields.io/badge/Arxiv-2411.19378-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2411.19378) \n[![GitHub star chart](https://img.shields.io/github/stars/X-iZhang/Libra?style=social)](https://star-history.com/#X-iZhang/Libra)\n[![License](https://img.shields.io/badge/License-Apache%202.0-yellow.svg?)](https://github.com/X-iZhang/Libra/blob/main/LICENSE)\n[![Visitors](https://api.visitorbadge.io/api/combined?path=https%3A%2F%2Fgithub.com%2FX-iZhang%2FLibra\u0026label=Views\u0026countColor=%23f36f43\u0026style=flat)](https://visitorbadge.io/status?path=https%3A%2F%2Fgithub.com%2FX-iZhang%2FLibra)\n\u003c/div\u003e\n\n*This repository hosts **Libra**, a tool designed to generate radiology reports by leveraging temporal information from chest X-rays taken at different time points.*\n\n\u003cdetails open\u003e\u003csummary\u003e📢 More Than Radiology: Codespace Features for MLLMs Workflow You’ll Love! 🎉 \u003c/summary\u003e\u003cp\u003e\n\n\u003e  * **Support for LLaVA-Type, LLaMA 3, Mistral, Phi-3/4 \u0026 Gemma**: Effortlessly run and fine-tune a variety of advanced open models.\n\u003e  * **Resume Training**: Resume training from checkpoints at any stage, whether for pre-training or fine-tuning.  \n\u003e  * **Validation Dataset**: Track model performance in real-time on `validation datasets` during training. \n\u003e  * **Custom Metrics**: Go beyond `eval_loss` with metrics like `BLEU`, `ROUGE-L`, `RadGraph-F1` or define your own criteria on valid dataset.   \n\u003e  * **Smart Saving**: Automatically save the best model based on validation loss or custom evaluation scores.\n\n\u003c/p\u003e\u003c/details\u003e\n\n## 🔥 News\n- **[18 Jun 2025]** 🎤 Invited talk at [**HealTAC 2025**](https://healtac2025.github.io/programme/) — topic: [*Towards Temporal-Aware Multimodal Large Language Models for Improved Radiology Report Generation*](https://x-izhang.github.io/post/healtac2025/)\n- **[16 May 2025]** 📝 A short blog: some musings on [***\"What Does ‘Temporal’ Really Mean?”***](https://x-izhang.github.io/blog/libra-blog1/) — thoughts behind Libra and temporal reasoning in radiology.\n- **[15 May 2025]** 🥳 [***The paper***](https://arxiv.org/pdf/2411.19378v2) has been accepted to [**ACL 2025**](https://2025.aclweb.org/)!\n- **[09 May 2025]** ✨ Now with full support for the [Phi-4](https://huggingface.co/collections/microsoft/phi-4-677e9380e514feb5577a40e4) family — compact language and reasoning models from Microsoft.\n- **[24 Mar 2025]** 🏆 **Libra** was invited to the [**ReXrank**](https://rexrank.ai/) Challenge — a leading leaderboard for Chest X-ray Report Generation.\n- **[10 Mar 2025]**  ✅ The architecture of [LLaVA-Med v1.5](https://huggingface.co/microsoft/llava-med-v1.5-mistral-7b) is now supported by this repo. [**Compatible weights**](https://github.com/X-iZhang/Libra?tab=readme-ov-file#libra-v05) are provided, with 'unfreeze_mm_vision_tower: true' set to ensure the *adapted* vision encoder is used.\n- **[11 Feb 2025]** 🚨 [**Libra-v1.0-3b**](https://huggingface.co/X-iZhang/libra-v1.0-3b) has been released! A **Small Multimodal Language Model for Radiology Report Generation**,  following the same training strategy as **Libra**.\n\n\u003cdetails\u003e\n\u003csummary\u003e- More -\u003c/summary\u003e\n\n- **[10 Feb 2025]** 🚀 The [**Libra**](https://github.com/X-iZhang/Libra) repo now supports [Mistral](https://huggingface.co/mistralai), [Phi-3](https://huggingface.co/collections/microsoft/phi-3-6626e15e9585a200d2d761e3), and [Gemma](https://huggingface.co/collections/google/gemma-2-release-667d6600fd5220e7b967f315) as LLMs, along with [SigLip](https://huggingface.co/collections/google/siglip-659d5e62f0ae1a57ae0e83ba) as the encoder!\n- **[19 Jan 2025]** ⚡ The **online demo** is available at [Hugging Face Demo](https://huggingface.co/spaces/X-iZhang/Libra). Welcome to try it out!\n- **[07 Jan 2025]** 🗂️ The processed data is available at [Data Download](https://github.com/X-iZhang/Libra#data-download).\n- **[20 Dec 2024]** 🚨 [**Libra-v1.0-7b**](https://huggingface.co/X-iZhang/libra-v1.0-7b) has been released!\n\n\u003c/details\u003e\n\n\n## Overview\nRadiology report generation requires integrating temporal medical images and creating accurate reports. Traditional methods often overlook crucial temporal information. We introduce **Libra**, a temporal-aware MLLM for chest X-ray report generation. Libra combines a radiology-specific image encoder with a novel **`Temporal Alignment Connector (TAC)`**, designed to accurately capture and integrate temporal differences between paired current and prior images. Experiments show that Libra sets new performance benchmarks on the MIMIC-CXR dataset for the RRG task.\n\n\u003cdetails\u003e\n\u003csummary\u003eLibra’s Architecture\u003c/summary\u003e\n\n![architecture](./assets/libra_architecture.png)\n\n\u003c/details\u003e\n\n## Contents\n- [Install](#install)\n- [Model Weights](#model-weights)\n    - [Libra-v1.0](#libra-v10)\n    - [Libra-v0.5](#libra-v05)\n    - [Projector weights](#projector-weights)\n- [Quick Start](#quick-start)\n    - [Gradio Web UI](#gradio-web-ui)\n    - [CLI Inference](#cli-inference)\n    - [Script Inference](#script-inference)\n- [Dataset](#dataset)\n    - [Prepare Data](#prepare-data)\n    - [Preprocess Data](#preprocess-data)\n    - [Data Download](#data-download)\n- [Train](#train)\n    - [Hyperparameters](#hyperparameters)\n    - [Stage 1: visual feature alignment](#stage-1-visual-feature-alignment)\n    - [Stage 2: RRG downstream task fine-tuning](#stage-2-rrg-downstream-task-fine-tuning)\n    - [✨New Options to Note](#new-options-to-note)\n- [Evaluation](#evaluation)\n    - [Generate model responses](#1-generate-libra-responses)\n    - [Evaluate the generated report](#2-evaluate-the-generated-report)\n    - [Metrics](#metrics)\n\n## Install\nWe strongly recommend that you create an environment from scratch as follows:\n1. Clone this repository and navigate to Libra folder\n```bash\ngit clone https://github.com/X-iZhang/Libra.git\ncd Libra\n```\n\n2. Install Package\n```Shell\nconda create -n libra python=3.10 -y\nconda activate libra\npip install --upgrade pip  # enable PEP 660 support\npip install -e .\n```\n\n3. Install additional packages for Training and Evaluation cases\n```Shell\npip install -e \".[train,eval]\"\npip install flash-attn --no-build-isolation\n```\n\n\u003cdetails\u003e\n\u003csummary\u003e Upgrade to latest code base \u003c/summary\u003e\n\n```Shell\ngit pull\npip install -e .\n```\n\n\u003c/details\u003e\n\n## Model Weights\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"./assets/result_chart.png\" width=\"350px\"\u003e \u003cbr\u003e\n  Libra-v1.0 achieves SoTA performance.\n\u003c/p\u003e\n\n### Libra-v1.0\n| Version | Size | Projector | Base LLM | Vision Encoder| Checkpoint |\n| ----------- | ----------- | ----------- | ----------- | ----------- | ----------- |\n| Libra-1.0 | 7B | TAC | Meditron-7B | RAD-DINO | [libra-v1.0-7b](https://huggingface.co/X-iZhang/libra-v1.0-7b) |\n| Libra-1.0 | 3B | TAC | Llama-3.2-3B-Instruct| RAD-DINO | [libra-v1.0-3b](https://huggingface.co/X-iZhang/libra-v1.0-3b) |\n\n\n\u003cdetails\u003e\n\u003csummary\u003e Performance on MIMIC-CXR (Findings section) \u003c/summary\u003e\n\n| Model | BLEU1 | BLEU4 | METEOR | ROUGE-L | RaTEScore | RG_ER |\n|----------|----------|-----------|-----------|---|---|---|\n| Libra-v1.0-7b | 51.3| 24.5 | 48.9 | 36.7 | 61.5 | 37.6 |\n| Libra-v1.0-3b | 50.5 | 23.3 | 48.5 | 35.2 | 61.1 | 37.5 |\n\n\u003c/details\u003e\n\n### Libra-v0.5\n\n| Version | Size | Projector | Base LLM | Vision Encoder| Checkpoint |\n| ----------- | ----------- | ----------- | ----------- | ----------- | ----------- |\n| Libra-0.5 | 7B | MLP-2x | Vicuna-7B | CLIP-L-336px | [Med-CXRGen-F](https://huggingface.co/X-iZhang/Med-CXRGen-F) |\n| Libra-0.5 | 7B | MLP-2x | Vicuna-7B | CLIP-L-336px | [Med-CXRGen-I](https://huggingface.co/X-iZhang/Med-CXRGen-I) |\n| Llava-med | 7B | MLP-2x | Mistral-7B-Instruct-v0.2 | CLIP-L-336px (adapted) | [Llava-Med-v1.5](https://huggingface.co/X-iZhang/libra-llava-med-v1.5-mistral-7b) |\n\n*💡Note: These two models are fine-tuned for `Findings` and `Impression` section generation. For more information on training strategies and dataset collection, please refer to [Med-CXRGen (Gla-AI4BioMed at RRG24)](https://github.com/X-iZhang/RRG-BioNLP-ACL2024)*\n\n### Projector weights\n\nThese projector weights were pre-trained for visual instruction tuning on chest X-ray to text generation tasks. They can be directly used to initialise your model for multimodal fine-tuning in similar clinical domains.\n\n⚠️ Important Note: For compatibility, please ensure that the *projector type*, *base LLM*, *conv_mode*, and *vision encoder* exactly match those used in our projector pretraining setup. Please also ensure the following settings are correctly configured during instruction tuning:\n\n```Shell\n--mm_projector_type TAC \\ # or mlp2x_gelu\n--mm_vision_select_layer all \\ # or -2\n--mm_vision_select_feature patch \\\n--mm_use_im_start_end False \\\n--mm_use_im_patch_token False \\\n```\n\n| Base LLM | conv_mode | Vision Encoder | Projector | Pretrain Data | Download |\n| ----------- | ----------- | ----------- | ----------- | ----------- | ----------- |\n| Meditron-7B | libra_v1 | RAD-DINO | TAC | [RRG \u0026 VQA](https://physionet.org/content/mimic-cxr) | [projector](https://huggingface.co/X-iZhang/libra-v1.0-7b/resolve/main/mm_tac_projector.bin) |\n| Llama-3.2-3B-Instruct | libra_llama_3 | RAD-DINO | TAC | [RRG \u0026 VQA](https://physionet.org/content/mimic-cxr) | [projector](https://huggingface.co/X-iZhang/libra-v1.0-3b/resolve/main/mm_tac_projector.bin) |\n| Vicuna-7B| libra_v0 | CLIP-L-336px| MLP-2x | [Findings section](https://huggingface.co/datasets/StanfordAIMI/rrg24-shared-task-bionlp) | [projector](https://huggingface.co/X-iZhang/Med-CXRGen-F/resolve/main/mm_mlp2x_projector_findings.bin) |\n| Vicuna-7B | libra_v0 | CLIP-L-336px | MLP-2x | [Impression section](https://huggingface.co/datasets/StanfordAIMI/rrg24-shared-task-bionlp) | [projector](https://huggingface.co/X-iZhang/Med-CXRGen-I/resolve/main/mm_mlp2x_projector_impressions.bin) |\n| Mistral-7B-Instruct-v0.2| llava_med_v1.5_mistral_7b | CLIP-L-336px | MLP-2x | [LLaVA-Med Dataset](https://github.com/microsoft/LLaVA-Med?tab=readme-ov-file#llava-med-dataset) | [projector](https://huggingface.co/X-iZhang/libra-llava-med-v1.5-mistral-7b/resolve/main/mm_mlp2x_projector_llavamed.bin?download=true) |\n\n## Quick Start\n\n### Gradio Web UI\n\nLaunch a local or online web demo by running:\n\n```bash\npython -m libra.serve.app\n```\n\n\u003cdetails\u003e\n\u003csummary\u003eSpecify your model:\u003c/summary\u003e\n\n```bash\npython -m libra.serve.app --model-path /path/to/your/model\n```\n\u003c/details\u003e\n\nYou just launched the Gradio web interface. Now, you can open the web interface with the URL printed on the screen. You will notice that both the default `libra-v1.0` model and `your model` are available in the model list, and you can choose to switch between them.\n\n![demo](./assets/demo.gif)\n\n### CLI Inference\nWe support running inference using the CLI. To use our model, run:\n```Shell\npython -m libra.serve.cli \\\n    --model-path X-iZhang/libra-v1.0-7b \\\n    --image-file \"./path/to/current_image.jpg\" \"./path/to/previous_image.jpg\"\n    # If there is no previous image, only one path is needed.\n```\n\n### Script Inference\nYou can use the `libra_eval` function in `libra/eval/run_libra.py` to easily launch a model trained by yourself or us on local machine or in Google Colab, after installing this repository.\n\n```Python\nfrom libra.eval import libra_eval\n\n# Define the model path, which can be a pre-trained model or your own fine-tuned model.\nmodel_path = \"X-iZhang/libra-v1.0-7b\"  # Or your own model\n\n# Define the paths to the images. The second image is optional for temporal comparisons.\nimage_files = [\n    \"./path/to/current/image.jpg\", \n    \"./path/to/previous/image.jpg\"  # Optional: Only include if a reference image is available\n]\n\n# Define the prompt to guide the model's response. Add clinical instructions if needed.\nprompt = (\n    \"Provide a detailed description of the findings in the radiology image. \"\n    \"Following clinical context: ...\"\n)\n\n# Specify the conversational mode, matching the PROMPT_VERSION used during training.\nconv_mode = \"libra_v1\"\n\n# Call the libra_eval function.\nlibra_eval(\n    model_path=model_path,\n    image_file=image_files,\n    query=prompt,\n    temperature=0.9,\n    top_p=0.8,\n    conv_mode=conv_mode,\n    max_new_tokens=512\n)\n```\n\u003cdetails\u003e\n\u003csummary\u003eMeanwhile, you can use the Beam Search method to obtain output.\u003c/summary\u003e\n\n```Python\nlibra_eval(\n    model_path=model_path,\n    image_file=image_files,\n    query=prompt,\n    num_beams=5, \n    length_penalty=2,\n    num_return_sequences=2,\n    conv_mode=conv_mode,\n    max_new_tokens=512\n)\n```\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eAdditionally, you can directly use LoRA weights for inference.\u003c/summary\u003e\n\n```Python\nlibra_eval(\n    model_path=\"./path/to/lora_weights\",  # path to LoRA weights\n    model_base=\"./path/to/base_model\",  # path to base Libra model\n    image_file=image_files,\n    query=prompt,\n    num_beams=5, \n    length_penalty=2,\n    num_return_sequences=2,\n    conv_mode=conv_mode,\n    max_new_tokens=512\n)\n```\n\n\u003c/details\u003e\n\n## Dataset\n\n### Prepare Data\n\nAll the data we use comes from [MIMIC-CXR](https://physionet.org/content/mimic-cxr/2.0.0/) and its two variants, and we strictly follow the official split for `train/valid/test` division.\n\n- Image Data\n\nAll images used for **Libra** come from the [MIMIC-CXR-JPG](https://physionet.org/content/mimic-cxr-jpg/2.0.0/) dataset in `.jpg` format. `DICOM` format is also supported and can be found in the [MIMIC-CXR](https://physionet.org/content/mimic-cxr/2.0.0/).\n\nAfter downloading the images, they will be automatically organized into the following structure in `./path/to/playground/data`:\n\n```\n./data/physionet.org/files/mimic-cxr-jpg/2.0.0\n└──files\n    ├── p10\n    │   └── p10000032\n    │       └── s50414267\n    │           ├── image1.jpg\n    │           └── image2.jpg\n    ├── p11\n    ├── p12\n    ├── ...\n    └── p19\n```\n\n- Annotation Data\n\nAll annotations used for **Libra** come from the [MIMIC-CXR](https://physionet.org/content/mimic-cxr/2.0.0/) and its two variants. This includes Radiology Reports and other relevant Visual Question Answering. \n\nPlease download the following datasets from the official website: `mimic-cxr-reports.zip` from [MIMIC-CXR](https://physionet.org/content/mimic-cxr/2.0.0/), [MIMIC-Diff-VQA](https://physionet.org/content/medical-diff-vqa/1.0.0/), and [MIMIC-Ext-*MIMIC-CXR-VQA*](https://physionet.org/content/mimic-ext-mimic-cxr-vqa/1.0.0/).\n\n### Preprocess Data\n\n- Radiology Report Sections\n\nFor free-text radiology report, we extract the `Findings`, `Impression`, `Indication`, `History`, `Comparison`, and `Technique` sections using the official [mimic-cxr](https://github.com/MIT-LCP/mimic-cxr/tree/master/txt) repository.\n\n*💡Note: To enable more structured and accurate extraction of the `Indication`, `History`, `Comparison`, and `Technique` sections—beyond what the original scripts provide—we replace the official ` .py` with our customised versions located in [`Libra/scripts/mimic-cxr/`](./scripts/mimic-cxr/).*\n\n\n- Visual Question Answering for Chest X-ray\n\nIn [Medical-Diff-VQA](https://physionet.org/content/medical-diff-vqa/1.0.0/), the main image is used as the current image, and the reference image is used as the prior image. In [MIMIC-Ext-MIMIC-CXR-VQA](https://physionet.org/content/mimic-ext-mimic-cxr-vqa/1.0.0/), all cases use a dummy prior image.\n\n### Data Download\n\n| Alignment data files | Split | Size |\n| ----- | ----- | -----: |\n| [libra_alignment_train.json](https://drive.google.com/file/d/1AIT1b3eRXgJFp3FJmHci3haTunK1NTMA/view?usp=drive_link)| train | 780 MiB |\n| [libra_alignment_valid.json](https://drive.google.com/file/d/1nvbUoDmw7j4HgXwZWiiACIhvZ6BvR2LX/view?usp=sharing)| valid | 79 MiB |\n\n| Fine-Tuning data files | Split | Size |\n| ----- | ----- | ----- |\n| [libra_findings_section_train.json](https://drive.google.com/file/d/1rJ3G4uiHlzK_P6ZBUbAi-cDaWV-o6fcz/view?usp=sharing)| train | 159 MiB |\n| [libra_findings_section_valid.json](https://drive.google.com/file/d/1IYwQS23veOU5SXWGYiTyq9VHUwkVESfD/view?usp=sharing)| valid | 79 MiB |\n\n| Evaluation data files | Split | Size |\n| --- | --- | ---: |\n| [libra_findings_section_eval.jsonl](https://drive.google.com/file/d/1fy_WX616L8SgyAonadJ2fUIEaX0yrGrQ/view?usp=sharing)| eval | 2 MiB |\n\n\n\u003cdetails\u003e\n\n\u003csummary\u003eMeanwhile, here are some bonus evaluation data files.\u003c/summary\u003e\n\n| Evaluation data files | Split | Size |\n| --- | --- | ---: |\n| [libra_impressions_section_eval.jsonl](https://drive.google.com/file/d/16msRfk7XxCmq7ZPG82lKvsnnjqsRPv__/view?usp=sharing)| eval | 1 MiB |\n| [libra_MIMIC-Ext-MIMIC-CXR-VQA_eval.jsonl](https://drive.google.com/file/d/1krPMwGGY6HP4sonNKlnkhLOoZrdjfVMW/view?usp=sharing)| eval | 4 MiB |\n| [libra_MIMIC-Diff-VQA _eval.jsonl](https://drive.google.com/file/d/1tP_CxPMM9PiKTq1mLYRHICcyJ36Q13mC/view?usp=sharing)| eval | 20 MiB |\n\n\u003c/details\u003e\n\n\n\nIf you want to train or evaluate your own tasks or datasets, please refer to [`Custom_Data.md`](https://github.com/X-iZhang/Libra/blob/main/CUSTOM_DATA.md).\n\n\n## Train\nLibra adopt a two-stage training strategy: (1) visual feature alignment: the visual encoder and LLM weights are frozen, and the Temporal Alignment Connector is trained; (2) RRG downstream task fine-tuning: apply LoRA to fine-tune the pre-trained LLM on the Findings section generation task.\n\nLibra is trained on 1 A6000 GPU with 48GB memory. To train on multiple GPUs, you can set the `per_device_train_batch_size` and the `gradient_accumulation_steps` accordingly. Always keep the global batch size the same: `per_device_train_batch_size` x `gradient_accumulation_steps` x `num_gpus`.\n\n### Hyperparameters\nWe set reasonable hyperparameters based on our device. The hyperparameters used in both pretraining and LoRA finetuning are provided below.\n\n1. Pretraining\n\n| Hyperparameter | Global Batch Size | Learning rate | Epochs | Max length | Weight decay |\n| --- | :---: | :---: | :---: | :---: | :---: |\n| Libra-v1.0-7b | 16 | 2e-5 | 1 | 2048 | 0 |\n\n2. LoRA finetuning\n\n| Hyperparameter | Global Batch Size | Learning rate | Epochs | Max length | Weight decay | LoRA rank | LoRA alpha |\n| --- | :---: | :---: | :---: | :---: | :---: | :---: | :---: |\n| Libra-v1.0-7b | 16 | 2e-5 | 3 | 2048 | 0 | 128 | 256 |\n\n### Download Meditron checkpoints (automatically)\n\nOur base LLM model, [Meditron-7B](https://huggingface.co/epfl-llm/meditron-7b), adapted to the medical domain from the Llama-2-7B model, will be downloaded automatically when you run our provided training scripts. No action is needed on your part.\n\n### Stage 1: visual feature alignment\n\nPretraining takes approximately 385 hours for Libra-v1.0-7b-pretrain on a single A6000 GPU (48GB) due to device limitations.\n\nFor detailed training scripts and guidelines, please refer to the following: [`pretrain.sh`](https://github.com/X-iZhang/Libra/blob/main/scripts/pretrain.sh) and [`pretrain_xformers.sh`](https://github.com/X-iZhang/Libra/blob/main/scripts/pretrain_xformers.sh) for [memory-efficient attention](https://arxiv.org/abs/2112.05682) implemented in [xFormers](https://github.com/facebookresearch/xformers).\n\n- `--mm_projector_type TAC`: the Temporal Alignment Connector.\n- `--vision_tower microsoft/rad-dino`: RAD-DINO is a vision transformer for encoding chest X-rays using DINOv2.\n- `--mm_vision_select_layer all`: Use all image features from the encoder for the Layerwise Feature Extractor.\n- `--tune_mm_mlp_adapter True`\n- `--freeze_mm_mlp_adapter False` \n\n### Stage 2: RRG downstream task fine-tuning\nYou may download our pretrained projectors from the [`mm_tac_projector.bin`](https://huggingface.co/X-iZhang/libra-v1.0-7b) file. It takes around 213 hours for Libra-v1.0-7b on a single A6000 GPU (48GB) due to device limitations.\n\nFor detailed training scripts and guidelines, please refer to: [`finetune_lora.sh`](https://github.com/X-iZhang/Libra/blob/main/scripts/finetune_lora.sh).\n\n- `--tune_mm_mlp_adapter False`\n- `--freeze_mm_mlp_adapter True` \n\nIf you have enough GPU memory: Use [`finetune.sh`](https://github.com/X-iZhang/Libra/blob/main/scripts/finetune.sh) to fine-tune the entire model. Alternatively, you can replace `zero3.json` with `zero3_offload.json` to offload some parameters to CPU RAM, though this will slow down the training speed.\n\nIf you are interested in continue finetuning Libra model to your own task/data, please check out [`Custom_Data.md`](https://github.com/X-iZhang/Libra/blob/main/CUSTOM_DATA.md).\n\n### New Options to Note\n\n- `--mm_projector_type TAC`: Specifies the Temporal Alignment Connector for Libra.\n- `--vision_tower microsoft/rad-dino`: Uses RAD-DINO as the chest X-rays encoder.\n- `--mm_vision_select_layer all`: Selects specific vision layers (e.g., -1, -2) or \"all\" for all layers.\n- `--validation_data_path ./path/`: Path to the validation data.\n- `--compute_metrics True`: Optionally computes metrics during validation. Note that this can consume significant memory. If GPU memory is insufficient, it is recommended to either disable this option or use a smaller validation dataset.\n\n\n## Evaluation\nIn Libra-v1.0, we evaluate models on the MIMIC-CXR test split for the findings section generation task. You can download the evaluation data [here](https://drive.google.com/file/d/1fy_WX616L8SgyAonadJ2fUIEaX0yrGrQ/view?usp=sharing). To ensure reproducibility and output quality, we evaluate our model using the beam search strategy.\n\n### 1. Generate Libra responses.\n\n```Shell\npython -m libra.eval.eval_vqa_libra \\\n    --model-path X-iZhang/libra-v1.0-7b \\\n    --question-file libra_findings_section_eval.jsonl \\\n    --image-folder ./physionet.org/files/mimic-cxr-jpg/2.0.0 \\\n    --answers-file /path/to/answer-file.jsonl \\\n    --num_beams 10 \\\n    --length_penalty 2 \\\n    --max_new_tokens 1024 \\\n    --conv-mode libra_v1\n```\n\nYou can evaluate Libra on your custom datasets by converting your dataset to the [JSONL format](https://github.com/X-iZhang/Libra/blob/main/CUSTOM_DATA.md#evaluation-dataset-format) and evaluating using [`eval_vqa_libra.py`](https://github.com/X-iZhang/Libra/blob/main/libra/eval/eval_vqa_libra.py).\n\nAdditionally, you can execute the evaluation using the command line. For detailed instructions, see [`libra_eval.sh`](https://github.com/X-iZhang/Libra/blob/main/scripts/eval/libra_eval.sh).\n\n```bash\nbash ./scripts/eval/libra_eval.sh beam\n```\n\n### 2. Evaluate the generated report.\n\nIn our case, you can directly use `libra_findings_section_eval.jsonl` and `answer-file.jsonl` for basic evaluation, using [`radiology_report.py`](https://github.com/X-iZhang/Libra/blob/main/libra/eval/radiology_report.py).\n\n```Python\nfrom libra.eval import evaluate_report\n\nreferences = \"libra_findings_section_eval.jsonl\"\npredictions = \"answer-file.jsonl\"\n\nresul = evaluate_report(references=references, predictions=predictions)\n\n# Evaluation scores\nresul\n{'BLEU1': 51.25,\n 'BLEU2': 37.48,\n 'BLEU3': 29.56,\n 'BLEU4': 24.54,\n 'METEOR': 48.90,\n 'ROUGE-L': 36.66,\n 'Bert_score': 62.50,\n 'Temporal_entity_score': 35.34}\n```\nOr use the command line to evaluate multiple references and store the results in a `.csv` file. For detailed instructions, see [`get_eval_scores.sh`](https://github.com/X-iZhang/Libra/blob/main/scripts/eval/get_eval_scores.sh).\n\n```bash\nbash ./scripts/eval/get_eval_scores.sh\n```\n### Metrics\n- Temporal Entity F1\n\nThe $F1_{temp}$ score includes common radiology-related keywords associated with temporal changes. You can use [`temporal_f1.py`](https://github.com/X-iZhang/Libra/blob/main/libra/eval/temporal_f1.py) as follows:\n\n```Python\nfrom libra.eval import temporal_f1_score\n\npredictions = [\n    \"The pleural effusion has progressively worsened since previous scan.\",\n    \"The pleural effusion is noted again on the current scan.\"\n]\nreferences = [\n    \"Compare with prior scan, pleural effusion has worsened.\",\n    \"Pleural effusion has worsened.\"\n]\n\ntem_f1_score = temporal_f1_score(\n    predictions=predictions,\n    references=references\n)\n\n# Temporal Entity F1 score\ntem_f1_score\n{'f1': 0.500000000075,\n 'prediction_entities': [{'worsened'}, set()],\n 'reference_entities': [{'worsened'}, {'worsened'}]}\n```\n\n- Radiology-specific Metrics\n\nSome specific metrics may require configurations that could conflict with Libra. It is recommended to follow the official guidelines and use separate environments for evaluation: [`RG_ER`](https://pypi.org/project/radgraph/0.1.13/), [`CheXpert-F1`](https://pypi.org/project/f1chexbert/), [`RadGraph-F1, RadCliQ, CheXbert vector`](https://github.com/rajpurkarlab/CXR-Report-Metric).\n\n\u003c!-- ![architecture](./assets/libra_architecture.png) --\u003e\n\n## Acknowledgements 🙏\n\nWe sincerely thank the following projects for their contributions to **Libra**:\n\n* [LLaVA](https://github.com/haotian-liu/LLaVA): A Large Language and Vision Assistant, laying the groundwork for multimodal understanding.\n* [FastChat](https://github.com/lm-sys/FastChat): An Open Platform for Training, Serving, and Evaluating Large Language Model based Chatbots.\n* [LLaMA](https://github.com/facebookresearch/llama): Open and efficient foundation language models that inspired our core language processing capabilities.\n* [MEDITRON](https://github.com/epfLLM/meditron): Open and efficient medical Large language models.\n* [RAD-DINO](https://huggingface.co/microsoft/rad-dino): An open and efficient biomedical image encoder, enabling robust radiological analysis.\n\n## Citation ✒️\n\nIf you find our paper and code useful in your research and applications, please cite using this BibTeX:\n```BibTeX\n@misc{zhang2025libraleveragingtemporalimages,\n      title={Libra: Leveraging Temporal Images for Biomedical Radiology Analysis}, \n      author={Xi Zhang and Zaiqiao Meng and Jake Lever and Edmond S. L. Ho},\n      year={2025},\n      eprint={2411.19378},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2411.19378}, \n}\n```\n## Intended Use 🧰\n\nLibra is primarily designed to **assist** clinical practitioners, researchers, and medical students in generating chest X-ray reports. Key applications include:\n\n- **Clinical Decision Support**: Providing draft findings that can be refined by a radiologist.  \n- **Educational Tool**: Demonstrating example interpretations and temporal changes for training radiology residents.  \n- **Research**: Facilitating studies on automated report generation and temporal feature learning in medical imaging.\n\n\u003e **Important**: Outputs should be reviewed by qualified radiologists or medical professionals before final clinical decisions are made.\n\n\u003cdetails\u003e\n\u003csummary\u003eLimitations and Recommendations\u003c/summary\u003e\n\n1. **Data Bias**: The model’s performance may be less reliable for underrepresented demographics or rare pathologies.  \n2. **Clinical Oversight**: Always involve a medical professional to verify the results—Libra is not a substitute for professional judgment.  \n3. **Temporal Inaccuracies**: Despite TAC’s focus on temporal alignment, subtle or uncommon changes may go unrecognized.  \n4. **Generalization**: Libra’s performance on chest X-ray types or conditions not seen during training may be limited.\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eEthical Considerations\u003c/summary\u003e\n\n- **Patient Privacy**: Ensure the data is fully de-identified and compliant with HIPAA/GDPR (or relevant privacy regulations).  \n- **Responsible Use**: Deploy Libra’s outputs carefully; they are not guaranteed to be error-free.  \n- **Accountability**: Users and organizations must assume responsibility for verifying clinical accuracy and safety.\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eDisclaimer\u003c/summary\u003e\n\nThis tool is for research and educational purposes only. It is not FDA-approved or CE-marked for clinical use. Users should consult qualified healthcare professionals for any clinical decisions.\n\u003c/details\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fx-izhang%2Flibra","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fx-izhang%2Flibra","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fx-izhang%2Flibra/lists"}