{"id":19279783,"url":"https://github.com/showlab/gui-narrator","last_synced_at":"2025-06-21T01:35:53.006Z","repository":{"id":244960968,"uuid":"815936403","full_name":"showlab/GUI-Narrator","owner":"showlab","description":"Repository of GUI Action Narrator","archived":false,"fork":false,"pushed_at":"2025-04-08T09:13:32.000Z","size":20836,"stargazers_count":10,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-06-07T16:13:26.583Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"JavaScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/showlab.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-06-16T15:27:03.000Z","updated_at":"2025-04-11T07:41:45.000Z","dependencies_parsed_at":"2025-04-11T23:01:42.802Z","dependency_job_id":null,"html_url":"https://github.com/showlab/GUI-Narrator","commit_stats":null,"previous_names":["showlab/gui-narrator"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/showlab/GUI-Narrator","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/showlab%2FGUI-Narrator","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/showlab%2FGUI-Narrator/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/showlab%2FGUI-Narrator/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/showlab%2FGUI-Narrator/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/showlab","download_url":"https://codeload.github.com/showlab/GUI-Narrator/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/showlab%2FGUI-Narrator/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":261047096,"owners_count":23102381,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-09T21:16:06.775Z","updated_at":"2025-06-21T01:35:52.966Z","avatar_url":"https://github.com/showlab.png","language":"JavaScript","funding_links":[],"categories":[],"sub_categories":[],"readme":"## GUI Action Narrator: Where and When Did That Action Take Place?\r\n\r\nQinchen Wu, Difei Gao, Kevin Qinghong Lin, Zhuoyu Wu, Xiangwu Guo, Peiran Li, Weichen Zhang, Hengxu Wang, Mike Zheng Shou\r\n\r\n\u003c!-- [![Project Website](https://img.shields.io/badge/Project-Website-blue)](https://showlab.github.io/GUI-Narrator/) --\u003e\r\n\r\n## 🤖: Introduction\r\n\r\nWe introduce GUI action dataset **Act2Cap** as well as an effective framework: **GUI Narrator** for GUI video captioning that utilizes the cursor detection to enhance the interpretation of high-resolution screenshots and keyframe extraction in GUI actions.\r\n\r\n\r\n## 📋 ToDo List\r\n\r\n- [x] Model for Cursor detector and Narrator\r\n- [ ] Code of conduct\r\n\r\n\r\n-- Our model and test benchmark are availble on  [![Hugging Face](https://img.shields.io/badge/Demo-HuggingFace-blue)](https://huggingface.co/FRank62Wu/ShowUI-Narrator).\r\n\r\n\r\n\r\n\r\n\r\n\r\n\u003c!-- - Download **ACT2CAP** dataset, which consists of 10-frame GUI screenshot sequences depicting atomic actions. **[Download link here](https://drive.google.com/file/d/18cL3ByBkEMI-eTKrelaEXWeiF3QwZAAl/view?usp=drive_link)**.\r\n- Narrations based on 10 frames screenshots in `.data_annotation` . Please replace the  `\u003cpath\u003e`  placeholder with the root path of ACT2CAP image files\r\n    ```{\r\n    \"id\": \"identity_3\",\r\n    \"conversations\": [\r\n        {\r\n        \"from\": \"user\",\r\n        \"value\": \"Picture1: \u003cimg\u003e.\u003cpath\u003e/action_video_10_frames/x/a_prompt.png\u003c/img\u003e\\n\r\n        Picture2: \u003cimg\u003e.\u003cpath\u003e/action_video_10_frames/x/b_prompt.png\u003c/img\u003e\\n\r\n        Picture3: \u003cimg\u003e.\u003cpath\u003e/action_video_10_frames/x/a_crop.png\u003c/img\u003e\\n\r\n        Picture4: \u003cimg\u003e.\u003cpath\u003e/action_video_10_frames/x/b_crop.png\u003c/img\u003e\\n \r\n        the images shows video clips of an atomic action on graphic user interface. The cursor is acting in the green bounding box\\nDescribe what is the cursor doing based on the given images. Leftclick, Rightclick, Doubleclick, Type write or Drag.\"\r\n        },\r\n        {\r\n        \"from\": \"assistant\",\r\n        \"value\": \"The cursor LeftClick on Swap\"\r\n        }\r\n    ]\r\n    }\r\n    ```\r\n    Where `a`, `b` denotes the start and the end frame index respectively. `x` denotes the folder index.\r\n    The terms `Prompt` and `Crop` refers to screen shot with visual prompt and cropped detailed images generated depend on cursor detection module. \r\n    However, if you are interested in the original images, you can substitute them with `frame_idx`. \r\n---\r\n- Download **Cursor detection and Key frame extraction checkpoint** from **[Download link here](https://drive.google.com/file/d/1ChrpBuPL7W84mKNsSsbueff5EGlyB3h2/view?usp=sharing)**\r\n\r\n- Import supporting packages\r\n  ```\r\n  pip install -r requirements.txt\r\n  ```\r\n\r\n- Run inference code as below, the visual prompts and cropped images will be generated in folder `frames_sample `\r\n   ``` \r\n       cd model\r\n       python run_model.py \\\r\n       --frame_extract_model_path /path/to/checkpoint_key_frames \\\r\n       --yolo_model_path /path/to/Yolo_best \\\r\n       --images_path /path/to/frames_sample \r\n   ``` --\u003e","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fshowlab%2Fgui-narrator","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fshowlab%2Fgui-narrator","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fshowlab%2Fgui-narrator/lists"}