{"id":31909015,"url":"https://github.com/allenai/manipulathor","last_synced_at":"2025-10-13T16:00:25.407Z","repository":{"id":37051959,"uuid":"355720441","full_name":"allenai/manipulathor","owner":"allenai","description":"ManipulaTHOR, a framework that facilitates visual manipulation of objects using a robotic arm","archived":false,"fork":false,"pushed_at":"2023-02-07T19:09:04.000Z","size":118353,"stargazers_count":86,"open_issues_count":4,"forks_count":13,"subscribers_count":10,"default_branch":"main","last_synced_at":"2024-04-14T07:49:58.871Z","etag":null,"topics":["ai2thor-environment","embodied-ai","manipulathor","object-manipulation"],"latest_commit_sha":null,"homepage":"https://ai2thor.allenai.org/manipulathor","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/allenai.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2021-04-08T00:43:25.000Z","updated_at":"2024-03-30T09:35:04.000Z","dependencies_parsed_at":"2023-02-12T06:15:28.338Z","dependency_job_id":null,"html_url":"https://github.com/allenai/manipulathor","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/allenai/manipulathor","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fmanipulathor","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fmanipulathor/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fmanipulathor/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fmanipulathor/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/allenai","download_url":"https://codeload.github.com/allenai/manipulathor/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fmanipulathor/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279015953,"owners_count":26085777,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-13T02:00:06.723Z","response_time":61,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai2thor-environment","embodied-ai","manipulathor","object-manipulation"],"created_at":"2025-10-13T16:00:04.863Z","updated_at":"2025-10-13T16:00:25.385Z","avatar_url":"https://github.com/allenai.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# \u003ca href=\"https://arxiv.org/pdf/2104.11213.pdf\"\u003eManipulaTHOR: A Framework for Visual Object Manipulation\u003c/a\u003e\n#### Kiana Ehsani, Winson Han, Alvaro Herrasti, Eli VanderBilt, Luca Weihs, Eric Kolve, Aniruddha Kembhavi, Roozbeh Mottaghi\n#### (Oral Presentation at CVPR 2021)\n#### \u003ca href=\"https://prior.allenai.org/projects/manipulathor\"\u003e(Project Page)\u003c/a\u003e--\u003ca href=\"http://ai2thor.allenai.org/manipulathor\"\u003e(Framework)\u003c/a\u003e--\u003ca href=\"https://www.youtube.com/watch?v=nINZ52nlzX0\u0026ab_channel=AllenInstituteforAI\"\u003e(Video)\u003c/a\u003e--\u003ca href=\"#TODO\"\u003e(Slides)\u003c/a\u003e \n\nWe present \u003cb\u003eManipulaTHOR\u003c/b\u003e, a framework that facilitates \u003cb\u003evisual manipulation\u003c/b\u003e of objects using a robotic arm. Our framework is built upon a \u003cb\u003ephysics engine\u003c/b\u003e and enables \u003cb\u003erealistic interactions\u003c/b\u003e with objects while navigating through scenes and performing tasks. Object manipulation is an established research domain within the robotics community and poses several challenges including \u003cb\u003eavoiding collisions\u003c/b\u003e, \u003cb\u003egrasping\u003c/b\u003e, and \u003cb\u003elong-horizon planning\u003c/b\u003e. Our framework focuses primarily on manipulation in visually rich and \u003cb\u003ecomplex scenes\u003c/b\u003e, \u003cb\u003ejoint manipulation and navigation\u003c/b\u003e planning, and \u003cb\u003egeneralization\u003c/b\u003e to unseen environments and objects; challenges that are often overlooked. The framework provides a comprehensive suite of sensory information and motor functions enabling development of robust manipulation agents.\n\nThis code base is based on \u003ca href=https://allenact.org/\u003eAllenAct\u003c/a\u003e framework and the majority of the core training algorithms and pipelines are borrowed from \u003ca href=https://github.com/allenai/allenact\u003eAllenAct code base\u003c/a\u003e. \n\n### Citation\n\nIf you find this project useful in your research, please consider citing:\n\n```\n   @inproceedings{ehsani2021manipulathor,\n     title={ManipulaTHOR: A Framework for Visual Object Manipulation},\n     author={Ehsani, Kiana and Han, Winson and Herrasti, Alvaro and VanderBilt, Eli and Weihs, Luca and Kolve, Eric and Kembhavi, Aniruddha and Mottaghi, Roozbeh},\n     booktitle={CVPR},\n     year={2021}\n   }\n```\n\n### Contents\n\u003cdiv class=\"toc\"\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"#-installation\"\u003e💻 Installation\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-armpointnav-task-description\"\u003e📝 ArmPointNav Task Description\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-dataset\"\u003e📊 Dataset\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-sensory-observations\"\u003e🖼️ Sensory Observations\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-allowed-actions\"\u003e🏃 Allowed Actions\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-defining-a-new-task\"\u003e✨Defining a New Task\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-training-an-agent\"\u003e🏋 Training an Agent\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"#-evaluating-a-pre-trained-agent\"\u003e💪 Evaluating a Pre-Trained Agent\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/div\u003e\n\n## 💻 Installation\n \nTo begin, clone this repository locally\n```bash\ngit clone https://github.com/ehsanik/manipulathor.git\n```\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eSee here for a summary of the most important files/directories in this repository\u003c/b\u003e \u003c/summary\u003e \n\u003cp\u003e\n\nHere's a quick summary of the most important files/directories in this repository:\n* `manipulathor_utils/*.py` - Helper functions and classes.\n* `manipulathor_baselines/armpointnav_baselines`\n    - `experiments/`\n        + `ithor/armpointnav_*.py` - Different baselines introduced in the paper. Each files in this folder corresponds to a row of a table in the paper.\n        + `*.py` - The base configuration files which define experiment setup and hyperparameters for training.\n    - `models/*.py` - A collection of Actor-Critic baseline models.  \n* `ithor_arm/` - A collection of Environments, Task Samplers and Task Definitions\n    - `ithor_arm_environment.py` - The definition of the `ManipulaTHOREnvironment` that wraps the AI2THOR-based framework introduced in this work and enables an easy-to-use API.  \n    - `itho_arm_constants.py` - Constants used to define the task and parameters of the environment. These include the step size \n      taken by the agent, the unique id of the the THOR build we use, etc.\n    - `ithor_arm_sensors.py` - Sensors which provide observations to our agents during training. E.g. the\n      `RGBSensor` obtains RGB images from the environment and returns them for use by the agent. \n    - `ithor_arm_tasks.py` - Definition of the `ArmPointNav` task, the reward definition and the function for calculating the goal achievement. \n    - `ithor_arm_task_samplers.py` - Definition of the `ArmPointNavTaskSampler` samplers. Initializing the sampler, reading the json files from the dataset and randomly choosing a task is defined in this file. \n    - `ithor_arm_viz.py` - Utility functions for visualization and logging the outputs of the models.\n\n\u003c/p\u003e\n\u003c/details\u003e\n\nYou can then install requirements by running\n```bash\npip install -r requirements.txt\n```\n\n\n\n**Python 3.6+ 🐍.** Each of the actions supports `typing` within \u003cspan class=\"chillMono\"\u003ePython\u003c/span\u003e.\n\n**AI2-THOR \u003cbcc2e6\u003e 🧞.** To ensure reproducible results, please install this version of the AI2THOR.\n\nAfter installing the requirements, you should start the xserver by running [this script](scripts/startx.py) in the background. Finally, you can start playing with the environment using [our example jupyter notebook](scripts/example_notebook.ipynb).\n\n## 📝 ArmPointNav Task Description\n\n\u003cimg src=\"media/armpointnav_task.png\" alt=\"\" width=\"100%\"\u003e\n\nArmPointNav is the goal of addressing the problem of visual object manipulation, where the task is to move an object between two locations in a scene. Operating in visually rich and complex environments, generalizing to unseen environments and objects, avoiding collisions with objects and structures in the scene, and visual planning to reach the destination are among the major challenges of this task. The example illustrates a sequence of actions taken a by a virtual robot within the ManipulaTHOR environment for picking up a vase from the shelf and stack it on a plate on the countertop.\n   \n## 📊 Dataset\n\nTo study the task of ArmPointNav, we present the ArmPointNav Dataset (APND). This consists of 30 kitchen scenes in AI2-THOR that include more than 150 object categories (69 interactable object categories) with a variety of shapes, sizes and textures. We use 12 pickupable categories as our target objects. We use 20 scenes in the training set and the remaining is evenly split into Val and Test. We train with 6 object categories and use the remaining to test our model in a Novel-Obj setting. For more information on dataset, and how to download it refer to \u003ca href=\"datasets/README.md\"\u003eDataset Details\u003c/a\u003e.\n\n## 🖼️ Sensory Observations\n\nThe types of sensors provided for this paper include:\n\n1. **RGB images** - having shape `224x224x3` and an FOV of 90 degrees.  \n2. **Depth maps** - having shape `224x224` and an FOV of 90 degrees.\n3. **Perfect egomotion** - We allow for agents to know precisely what the object location is relative to the agent's arm as well as to its goal location.\n\n\n## 🏃 Allowed Actions\n\nA total of 13 actions are available to our agents, these include:\n\n1. **Moving the agent**\n\n* `MoveAhead` - Results in the agent moving ahead by 0.25m if doing so would not result in the agent colliding with something.\n\n* `Rotate [Right/Left]` - Results in the agent's body rotating 45 degrees by the desired direction.\n\n2. **Moving the arm**\n\n* `Moving the wrist along axis [x, y, z]` - Results in the arm moving along an axis (\u003cspan\u003e\u0026#177;\u003c/span\u003ex,\u003cspan\u003e\u0026#177;\u003c/span\u003ey, \u003cspan\u003e\u0026#177;\u003c/span\u003ez) by 0.05m.\n\n* `Moving the height of the arm base [Up/Down]` - Results in the base of the arm moving along y axis by 0.05m.\n\n3. **Abstract Grasp**\n\n* Picks up a target object. Only succeeds if the object is inside the arm grasper.\n  \n4. **Done Action**\n\n* This action finishes an episode. The agent must issue a `Done` action when it reaches the goal otherwise the episode considers as a failure.\n\n## ✨ Defining a New Task\n\nIn order to define a new task, redefine the rewarding, try a new model, or change the enviornment setup, checkout our tutorial on defining a new task \u003ca href=\"DefineTask.md\"\u003ehere\u003c/a\u003e.\n\n## 🏋 Training An Agent\n\nFor running experiments first you need to add the project directory to your python path. You can train a model with a specific experiment setup by running one of the experiments below:\n\n```\nallenact manipulathor_baselines/armpointnav_baselines/experiments/ithor/\u003cEXPERIMENT-NAME\u003e -o experiment_output -s 1\n```\n\nWhere `\u003cEXPERIMENT-NAME\u003e` can be one of the options below:\n\n```\narmpointnav_no_vision -- No Vision Baseline\narmpointnav_disjoint_depth -- Disjoint Model Ablation\narmpointnav_rgb -- Our RGB Experiment\narmpointnav_rgbdepth -- Our RGBD Experiment\narmpointnav_depth -- Our Depth Experiment\n``` \n\n\n## 💪 Evaluating A Pre-Trained Agent \n\nTo evaluate a pre-trained model, (for example to reproduce the numbers in the paper), you can add\n`-t test -c \u003cWEIGHT-ADDRESS\u003e` to the end of the command you ran for training. \n\nIn order to reproduce the numbers in the paper, you need to download the pretrained models from \n[here](https://drive.google.com/file/d/1wZi_IL5d7elXLkAb4jOixfY0M6-ZfkGM/view?usp=sharing) and extract them \nto pretrained_models. The full list of experiments and their corresponding trained weights can be found\n[here](pretrained_models/EvaluateModels.md).\n\n```\nallenact manipulathor_baselines/armpointnav_baselines/experiments/ithor/\u003cEXPERIMENT-NAME\u003e -o test_out -s 1 -t test -c \u003cWEIGHT-ADDRESS\u003e\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fallenai%2Fmanipulathor","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fallenai%2Fmanipulathor","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fallenai%2Fmanipulathor/lists"}