{"id":28653871,"url":"https://github.com/tiger-ai-lab/pixel-reasoner","last_synced_at":"2025-06-13T07:07:56.435Z","repository":{"id":294037247,"uuid":"985804909","full_name":"TIGER-AI-Lab/Pixel-Reasoner","owner":"TIGER-AI-Lab","description":"Pixel-Level Reasoning Model trained with RL","archived":false,"fork":false,"pushed_at":"2025-06-12T00:05:55.000Z","size":20050,"stargazers_count":125,"open_issues_count":1,"forks_count":2,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-06-12T00:34:02.954Z","etag":null,"topics":["llm","multimodal","reasoning"],"latest_commit_sha":null,"homepage":"https://tiger-ai-lab.github.io/Pixel-Reasoner/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/TIGER-AI-Lab.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-05-18T15:07:10.000Z","updated_at":"2025-06-12T00:05:58.000Z","dependencies_parsed_at":"2025-06-04T09:39:45.532Z","dependency_job_id":"be23e1a8-fed1-49fe-a37a-52ca34d3fd67","html_url":"https://github.com/TIGER-AI-Lab/Pixel-Reasoner","commit_stats":null,"previous_names":["tiger-ai-lab/pixel-reasoner"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/TIGER-AI-Lab/Pixel-Reasoner","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TIGER-AI-Lab%2FPixel-Reasoner","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TIGER-AI-Lab%2FPixel-Reasoner/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TIGER-AI-Lab%2FPixel-Reasoner/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TIGER-AI-Lab%2FPixel-Reasoner/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/TIGER-AI-Lab","download_url":"https://codeload.github.com/TIGER-AI-Lab/Pixel-Reasoner/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TIGER-AI-Lab%2FPixel-Reasoner/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":259599331,"owners_count":22882357,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["llm","multimodal","reasoning"],"created_at":"2025-06-13T07:07:55.918Z","updated_at":"2025-06-13T07:07:56.418Z","avatar_url":"https://github.com/TIGER-AI-Lab.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning\n\n\u003ca target=\"_blank\" href=\"https://arxiv.org/abs/2505.15966\"\u003e\n\u003cimg style=\"height:22pt\" src=\"https://img.shields.io/badge/-Paper-red?style=flat\u0026logo=arxiv\"\u003e\u003c/a\u003e\n\u003ca target=\"_blank\" href=\"#\"\u003e\n\u003cimg style=\"height:22pt\" src=\"https://img.shields.io/badge/-Code-green?style=flat\u0026logo=github\"\u003e\u003c/a\u003e\n\n\u003ca target=\"_blank\" href=\"https://tiger-ai-lab.github.io/Pixel-Reasoner/\"\u003e\n\u003cimg style=\"height:22pt\" src=\"https://img.shields.io/badge/-🌐%20Website-blue?style=flat\"\u003e\u003c/a\u003e\n\n\u003ca target=\"_blank\" href=\"https://huggingface.co/TIGER-Lab/PixelReasoner-RL-v1\"\u003e\n\u003cimg style=\"height:22pt\" src=\"https://img.shields.io/badge/-🤗%20Models-red?style=flat\"\u003e\u003c/a\u003e\n\n\u003ca target=\"_blank\" href=\"https://huggingface.co/collections/TIGER-Lab/pixel-reasoner-682fe96ea946d10dda60d24e\"\u003e\n\u003cimg style=\"height:22pt\" src=\"https://img.shields.io/badge/-🤗%20Dataset-blue?style=flat\"\u003e\u003c/a\u003e\n\n\u003ca target=\"_blank\" href=\"https://huggingface.co/spaces/TIGER-Lab/Pixel-Reasoner\"\u003e\n\u003cimg style=\"height:22pt\" src=\"https://img.shields.io/badge/-🤗%20Demo-yellow?style=flat\"\u003e\u003c/a\u003e\n\n\u003cbr\u003e\n\u003cspan\u003e\n\u003cb\u003eAuthors:\u003c/b\u003e Alex Su\u003csup\u003e*\u003c/sup\u003e\n\u003ca class=\"name\" target=\"_blank\" href=\"https://HaozheH3.github.io\"\u003eHaozhe Wang\u003csup\u003e*\u003c/sup\u003e\u003csup\u003e\u0026dagger;\u003c/sup\u003e\u003c/a\u003e, \n\u003ca class=\"name\" target=\"_blank\" href=\"https://cs.uwaterloo.ca/~w2ren/\"\u003eWeiming Ren\u003c/a\u003e, \n\u003ca class=\"name\" target=\"_blank\" href=\"https://cse.hkust.edu.hk/~flin/\"\u003eFangzhen Lin\u003c/a\u003e,\n\u003ca class=\"name\" target=\"_blank\" href=\"https://wenhuchen.github.io/\"\u003eWenhu Chen\u003csup\u003e\u0026Dagger;\u003c/sup\u003e\u003c/a\u003e\n\u003cbr\u003e\n\u003csup\u003e*\u003c/sup\u003eEqual Contribution. \n\u003csup\u003e\u0026dagger;\u003c/sup\u003eProject Lead. \n\u003csup\u003e\u0026Dagger;\u003c/sup\u003eCorrespondence.\n\u003c/span\u003e\n\n\n\n## 🔥News\n- [2025/5/25] We made really fun demos! You can now play with the [**online demo**](https://huggingface.co/spaces/TIGER-Lab/Pixel-Reasoner). Have fun!\n- [2025/5/22] We released models-v1. Now actively working on data and code release.\n\n\n## Overview\n![overview](./assets/teaser.png)\n\n\u003cdetails\u003e\u003csummary\u003eAbstract\u003c/summary\u003e \nChain-of-thought reasoning has significantly improved the performance of Large Language Models (LLMs) across various domains. However, this reasoning process has been confined exclusively to textual space, limiting its effectiveness in visually intensive tasks. To address this limitation, we introduce the concept of reasoning in the pixel-space. Within this novel framework, Vision-Language Models (VLMs) are equipped with a suite of visual reasoning operations, such as zoom-in and select-frame. These operations enable VLMs to directly inspect, interrogate, and infer from visual evidences, thereby enhancing reasoning fidelity for visual tasks.\nCultivating such pixel-space reasoning capabilities in VLMs presents notable challenges, including the model's initially imbalanced competence and its reluctance to adopt the newly introduced pixel-space operations. We address these challenges through a two-phase  training approach. The first phase employs instruction tuning on synthesized reasoning traces to familiarize the model with the novel visual operations. Following this, a reinforcement learning (RL) phase leverages a curiosity-driven reward scheme to balance exploration between pixel-space reasoning and textual reasoning. With these visual operations, VLMs can interact with complex visual inputs, such as information-rich images or videos to proactively gather necessary information. We demonstrate that this approach significantly improves VLM performance across diverse visual reasoning benchmarks. Our 7B model, \\model, achieves 84\\% on V* bench, 74\\% on TallyQA-Complex, and 84\\% on InfographicsVQA, marking the highest accuracy achieved by any open-source model to date. These results highlight the importance of pixel-space reasoning and the effectiveness of our framework.\n\u003c/details\u003e\n\n### Models\nPlease check the [TIGER-Lab/PixelReasoner-RL-v1](https://huggingface.co/TIGER-Lab/PixelReasoner-RL-v1) and [TIGER-Lab/PixelReasoner-WarmStart](https://huggingface.co/TIGER-Lab/PixelReasoner-WarmStart)\n\n## 🚀Quick Start\nWe proposed two-staged post-training. The instruction tuning is adapted from Open-R1. The Curiosity-Driven RL is adapted from VL-Rethinker.\n\n### Running Instruction Tuning\n\nFollow these steps to start the instruction tuning process:\n\n1. **Installation**\n   - Navigate to the `instruction_tuning` folder\n   - Follow the detailed setup guide in [installation instructions](instruction_tuning/install/install.md)\n\n2. **Configuration**\n   -configure model and data path in sft.sh\n\n3. **Launch Training**\n   ```bash\n   bash sft.sh\n   ```\n\n### Running Curiosity-Driven RL\nFirst prepare data. Run the following will get the training data prepared under `curiosity_driven_rl/data` folder. \n```\ndataname=PixelReasoner-RL-Data\nexport hf_user=TIGER-Lab\ncd onestep_evaluation\nbash prepare.sh ${dataname}\n```\n\nThen download model [TIGER-Lab/PixelReasoner-RL-v1](https://huggingface.co/TIGER-Lab/PixelReasoner-RL-v1).\n\nUnder `curiosity_driven_rl` folder, install the environment following [the installation instructions](curiosity_driven_rl/installation.md).\n\nSet the data path, model path, wandb keys (if you want to use it) in `curiosity_driven_rl/scripts/train_vlm_multi.sh`.\n\nRun the following.\n```bash\ncd curiosity_driven_rl\n\nexport temperature=1.0\nexport trainver=\"dataname\"\nexport testver=\"testdataname\"\nexport filter=True # filtering zero advantages\nexport algo=group # default for grpo\nexport lr=10\nexport MAX_PIXELS=4014080 # =[max_image_token]x28x28\nexport sys=vcot # system prompt version\nexport mode=train # [no_eval, eval_only, train]\nexport policy=/path/to/policy\nexport rbuffer=512 # replay buffer size\nexport bsz=256 # global train batch size\nexport evalsteps=4 \nexport nactor=4 # 4x8 for actor if with multinode, no effect with single node\nexport nvllm=32 # 4x8 for sampling if with multinode, no effect with single node\nexport tp=1 # vllm tp, 1 for 7B\nexport repeat=1 # data repeat\nexport nepoch=3 # data epoch\nexport logp_bsz=1 # must be 1\nexport maxlen=10000 # generate_max_len\nexport tagname=Train\n\nbash ./scripts/train_vlm_multi.sh\n```\n**Note**: the number of prompts into vLLM inference is controlled by `--eval_batch_size_pergpu` during evaluation, and `args.rollout_batch_size // strategy.world_size` during training. Must set `logp_bsz=1` or `--micro_rollout_batch_size=1` for computing logprobs because model.generate() suffers from feature mismatch when batchsize \u003e 1.\n\n### One-Step Evaluation\nEvaluation data can be found in [the HF Collection](https://huggingface.co/collections/JasperHaozhe/evaldata-pixelreasoner-6846868533a23e71a3055fe9).\n\n#### Image-Based Benchmarks\nLet's take the vstar evaluation as an example. The HF data path is `JasperHaozhe/VStar-EvalData-PixelReasoner`.\n\n**1. Prepare Data**\n```\ndataname=VStar-EvalData-PixelReasoner\ncd onestep_evaluation\nbash prepare.sh ${dataname}\n```\nThe bash script will download from HF, process the image paths, and move the data to `curiosity_driven_rl/data`. \n\nCheck the folder `curiosity_driven_rl/data`, you will know the downloaded parquet file is named as `vstar.parquet`. \n\n**2. Inference and Evaluation**\n\nInstall the openrlhf according to `curiosity_driven_rl/installation.md`.\n\nUnder `curiosity_driven_rl` folder. Set `benchmark=vstar`, `working_dir`, `policypath`, and `savefolder`,`tagname` for saving evaluation results. Run the following.\n```\nbenchmark=vstar\nexport working_dir=\"/path/to/curiosity_driven_rl\"\nexport policy=\"/path/to/policy\"\nexport savefolder=tooleval\nexport nvj_path=\"/path/to/nvidia/nvjitlink/lib\" # in case the system cannot fiind the nvjit library\n############\nexport sys=vcot # define the system prompt\nexport MIN_PIXELS=401408\nexport MAX_PIXELS=4014080 # define the image resolution\nexport eval_bsz=64 # vllm will processes this many queries \nexport tagname=eval_vstar_bestmodel\nexport testdata=\"${working_dir}/data/${benchmark}.parquet\"\nexport num_vllm=8\nexport num_gpus=8\nbash ${working_dir}/scripts/eval_vlm_new.sh\n```\n#### Video-Based Benchmarks\nFor the MVBench, we extracted the frames from videos and construct the eval data that fits into our evaluation. The data is available in `JasperHaozhe/MVBench-EvalData-PixelReasoner`\n\n**1. Prepare Data**\n```\ndataname=MVBench-EvalData-PixelReasoner\ncd onestep_evaluation\nbash prepare.sh ${dataname}\n```\n**2. Inference and Evaluation**\n\nUnder `curiosity_driven_rl` folder. Set `benchmark=mvbench`, `working_dir`, `policypath`, and `savefolder`,`tagname` for saving evaluation results. Run the following.\n```\nbenchmark=mvbench\nexport working_dir=\"/path/to/curiosity_driven_rl\"\nexport policy=\"/path/to/policy\"\nexport savefolder=tooleval\nexport nvj_path=\"/path/to/nvidia/nvjitlink/lib\" # in case the system cannot fiind the nvjit library\n############\nexport sys=vcot # define the system prompt\nexport MIN_PIXELS=401408\nexport MAX_PIXELS=4014080 # define the image resolution\nexport eval_bsz=64 # vllm will processes this many queries \nexport tagname=eval_vstar_bestmodel\nexport testdata=\"${working_dir}/data/${benchmark}.parquet\"\nexport num_vllm=8\nexport num_gpus=8\nbash ${working_dir}/scripts/eval_vlm_new.sh\n```\n\n#### Inference of Qwen2.5-VL-Instruct\nSet `sys=notool` for textual reasoning with Qwen2.5-VL-Instruct . `sys=vcot` can trigger zero-shot use of visual operations but may also induce unexpected behaviors.\n\n## Possible Exceptions\n**1. Exceeding Model Context Length**\n\nIf you do sampling (e.g., during training) and set a larger `MAX_PIXELS`, you could encounter the following:\n```\nValueError: The prompt (total length 10819) is too long to fit into the model (context length 10240). Make sure that `max_model_len` is no smaller than the number of text tokens plus multimodal tokens. For image inputs, the number of image tokens depends on the number of images, and possibly their aspect ratios as well.\n```\n\nThis stems from:\n1. max model length is too short, try adjust `generate_max_len + prompt_max_len`\n2. too many image tokens, which means there could be too many images or the image resolution is too large and takes up many image tokens. \nTo address this problem during training, you could set smaller `MAX_PIXELS`, or you could set the max number images during training, via `max_imgnum` in `curiosity_driven_rl/openrlhf/trainer/ppo_utils/experience_maker.py`.\n\n**2. transformers and vLLM version mismatch**\n```\nInput and cos/sin must have the same dtype, got torch.float32 and torch.bfloat32.\n```\nTry reinstall transformers: `pip install --force-reinstall git+https://github.com/huggingface/transformers.git@9985d06add07a4cc691dc54a7e34f54205c04d4` and update vLLM. \n\n**3. logp_bsz=1**\n```\nException: dump-info key solutions: {} should be {}\n```\nMust set `logp_bsz=1` or `--micro_rollout_batch_size=1` for computing logprobs because model.generate() suffers from feature mismatch when batchsize \u003e 1.\n\n**4. Reproduction Mismath**\n\nCorrect values of `MAX_PIXELS` and `MIN_PIXELS` are crucial for reproducing the results. \n1. make sure the env variables are correctly set\n2. make sure the vLLM engines correctly read the env variables\n\nWhen using ray parallelism, chances are that the env variables are only set on one machine but not all machines. To fix this, add ray env variables as follows:\n```\nRUNTIME_ENV_JSON=\"{\\\"env_vars\\\": {\\\"MAX_PIXELS\\\": \\\"$MAX_PIXELS\\\", \\\"MIN_PIXELS\\\": \\\"$MIN_PIXELS\\\"}}\"\n\nray job submit --address=\"http://127.0.0.1:8265\" \\\n--runtime-env-json=\"$RUNTIME_ENV_JSON\" \\\n```\n\nThanks [@LiqiangJing](https://github.com/LiqiangJing) for feedback!\n\n## Contact\nContact Haozhe (jasper.whz@outlook.com) for direct solution of any bugs in RL.\n\nContact Muze for SFT.\n\n\n\n## Citation\nIf you find this work useful, please give us a free cite:\n```bibtex\n@article{pixelreasoner,\n      title={Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning},\n      author = {Su, Alex and Wang, Haozhe and Ren, Weiming and Lin, Fangzhen and Chen, Wenhu},\n      journal={arXiv preprint arXiv:2505.15966},\n      year={2025}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftiger-ai-lab%2Fpixel-reasoner","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ftiger-ai-lab%2Fpixel-reasoner","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftiger-ai-lab%2Fpixel-reasoner/lists"}