{"id":22416564,"url":"https://github.com/mbodiai/embodied-agents","last_synced_at":"2025-09-25T23:56:37.372Z","repository":{"id":242097067,"uuid":"804600851","full_name":"mbodiai/embodied-agents","owner":"mbodiai","description":"Seamlessly integrate state-of-the-art transformer models into robotics stacks","archived":false,"fork":false,"pushed_at":"2025-09-14T20:26:46.000Z","size":79251,"stargazers_count":239,"open_issues_count":14,"forks_count":29,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-09-23T11:52:36.797Z","etag":null,"topics":["agents","artificial-intelligence","diffusion","embodied","embodied-agent","embodied-agents","generative-ai","large-language-models","llm","mbodi","mbodiai","multimodal","robotics","transformer","vision-language-model","vlm"],"latest_commit_sha":null,"homepage":"https://mbodi.ai","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/mbodiai.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-05-22T22:45:44.000Z","updated_at":"2025-09-19T04:52:33.000Z","dependencies_parsed_at":"2024-12-24T00:18:31.657Z","dependency_job_id":"96bc6eb0-aae1-48d9-8fe5-288b206cb3b3","html_url":"https://github.com/mbodiai/embodied-agents","commit_stats":null,"previous_names":["mbodiai/mbodied-agents","mbodiai/embodied-agents"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/mbodiai/embodied-agents","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbodiai%2Fembodied-agents","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbodiai%2Fembodied-agents/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbodiai%2Fembodied-agents/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbodiai%2Fembodied-agents/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/mbodiai","download_url":"https://codeload.github.com/mbodiai/embodied-agents/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbodiai%2Fembodied-agents/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":276599953,"owners_count":25671120,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-09-23T02:00:09.130Z","response_time":73,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agents","artificial-intelligence","diffusion","embodied","embodied-agent","embodied-agents","generative-ai","large-language-models","llm","mbodi","mbodiai","multimodal","robotics","transformer","vision-language-model","vlm"],"created_at":"2024-12-05T15:16:30.430Z","updated_at":"2025-09-25T23:56:37.317Z","avatar_url":"https://github.com/mbodiai.png","language":"Python","funding_links":[],"categories":["Methods"],"sub_categories":[],"readme":"\u003cdiv align=\"left\"\u003e\n \u003cimg src=\"assets/logo.png\" width=200;/\u003e\n  \u003cdiv\u003e\u0026nbsp;\u003c/div\u003e\n  \u003cdiv align=\"left\"\u003e\n    \u003csup\u003e\n      \u003ca href=\"https://api.mbodi.ai\"\u003e\n        \u003ci\u003e\u003cfont size=\"4\"\u003eBenchmark, Explore, and Send API Requests Now\u003c/font\u003e\u003c/i\u003e\n      \u003c/a\u003e\n    \u003c/sup\u003e\n  \u003c/div\u003e\n  \u003cdiv\u003e\u0026nbsp;\u003c/div\u003e\n\n[![license](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![PyPI - Python Version](https://img.shields.io/pypi/pyversions/mbodied)](https://pypi.org/project/mbodied/)\n[![PyPI](https://img.shields.io/pypi/v/mbodied)](https://pypi.org/project/mbodied)\n[![Downloads](https://static.pepy.tech/badge/mbodied)](https://pepy.tech/project/mbodied)\n[![MacOS](https://github.com/mbodiai/opensource/actions/workflows/macos.yml/badge.svg?branch=main)](https://github.com/mbodiai/opensource/actions/workflows/macos.yml)\n[![Ubuntu](https://github.com/mbodiai/opensource/actions/workflows/ubuntu.yml/badge.svg)](https://github.com/mbodiai/opensource/actions/workflows/ubuntu.yml)\n\n📖 **Docs**: [docs](https://api.mbodi.ai/docs)\n\n🚀 **Simple Robot Agent Example:** [\u003cimg align=\"center\" src=\"https://colab.research.google.com/assets/colab-badge.svg\" /\u003e](https://colab.research.google.com/drive/1KN0JohcjHX42wABBHe-CxXP-NWJjbZts?usp=sharing) \u003c/br\u003e\n💻 **Simulation Example with [SimplerEnv](https://github.com/simpler-env/SimplerEnv):** [\u003cimg align=\"center\" src=\"https://colab.research.google.com/assets/colab-badge.svg\" /\u003e](https://colab.research.google.com/drive/18oiuw1yTxO5x-eT7Z8qNyWtjyYd8cECI?usp=sharing) \u003c/br\u003e\n🤖 **Motor Agent using OpenVLA:** [\u003cimg align=\"center\" src=\"https://colab.research.google.com/assets/colab-badge.svg\" /\u003e](https://colab.research.google.com/drive/1flnMrqyepGOO8J9rE6rehzaLdZPsw6lX?usp=sharing)\u003c/br\u003e\n⏺️ **Record Dataset on a Robot:** [\u003cimg align=\"center\" src=\"https://colab.research.google.com/assets/colab-badge.svg\" /\u003e](https://colab.research.google.com/drive/15UuFbMUJGEjqJ_7I_b5EvKvLCKnAc8bB?usp=sharing)\u003c/br\u003e\n\n🫡 **Support, Discussion, and How-To's:** \u003c/br\u003e\n[![](https://dcbadge.limes.pink/api/server/BPQ7FEGxNb?theme=discord\u0026?logoColor=pink)](https://discord.gg/BPQ7FEGxNb)\n\n\u003c/div\u003e\n\n**Updates:**\n\n**April 18, 2025 — embodied-agents v1.5**\n\n- Updated sensory endpoints\n- Added support for Google Gemini as the language agent backend\n- Enabled tool calling for Language Agent (OpenAI)\n- Added Retrieval-Augmented Generation (RAG) functionalities — see [example](examples/6_robot_with_rag.py)\n\n**Aug 28, 2024, embodied-agents v1.2**\n\n- New [Doc site](https://api.mbodi.ai/docs) is up!\n- Added the features to record dataset on [robot](mbodied/robots/robot.py) natively.\n- Added multiple new Sensory Agents, e.g., [depth estimation](mbodied/agents/sense/depth_estimation_agent.py), [object detection](mbodied/agents/sense/object_detection_agent.py), [image segmentation](mbodied/agents/sense/segmentation_agent.py) with public [API endpoints](https://api.mbodi.ai/sense/) hosted. And a simple cli `mbodied` for trying them.\n- Added [Auto Agent](mbodied/agents/auto/auto_agent.py) for dynamic agents selection.\n\n**June 30, 2024, embodied-agents v1.0**:\n\n- Added Motor Agent supporting OpenVLA with free [API endpoint](https://api.mbodi.ai/community-models) hosted.\n- Added Sensory Agent supporting e.g., 3D object pose detection.\n- Improved automatic dataset recording.\n- Agent now can make remote act calls to API servers e.g., Gradio, vLLM.\n- Bug fixes and performance improvements have been made.\n- PyPI project is renamed to `mbodied`.\n\n# embodied agents\n\n**embodied agents** is a toolkit for integrating large multi-modal models into existing robot stacks with just a few lines of code. It provides consistency, reliability, scalability and is configurable to any observation and action space.\n\n\u003cimg src=\"assets/new_demo.gif\" alt=\"Demo GIF\" style=\"width: 550px;\"\u003e\n\n- [embodied agents](#embodied-agents)\n  - [Overview](#overview)\n    - [Motivation](#motivation)\n    - [Goals](#goals)\n    - [Limitations](#limitations)\n    - [Scope](#scope)\n    - [Features](#features)\n    - [Endpoints](#endpoints)\n    - [Support Matrix](#support-matrix)\n    - [Roadmap](#roadmap)\n  - [Installation](#installation)\n  - [Getting Started](#getting-started)\n    - [Customize a Motion to fit a robot's action space.](#customize-a-motion-to-fit-a-robots-action-space)\n    - [Run a robotics transformer model on a robot.](#run-a-robotics-transformer-model-on-a-robot)\n    - [Notebooks](#notebooks)\n  - [The Sample Class](#the-sample-class)\n  - [API Reference](#api-reference)\n  - [Directory Structure](#directory-structure)\n  - [Contributing](#contributing)\n\n## Overview\n\nThis repository is broken down into 3 main components: **Agents**, **Data**, and **Hardware**. Inspired by the efficiency of the central nervous system, each component is broken down into 3 meta-modalities: **Language**, **Motion**, and **Sense**. Each agent has an `act` method that can be overridden and satisfies:\n\n- **Language Agents** always return a string.\n- **Motor Agents** always return a `Motion`.\n- **Sensory Agents** always return a `SensorReading`.\n\nFor convenience, we also provide **AutoAgent** which dynamically initializes the right agent for the specified task. See [API Reference](#auto-agent) below for more.\n\nA call to `act` or `async_act` can perform local or remote inference synchronously or asynchronously. Remote execution can be performed with [Gradio](https://www.gradio.app/docs/python-client/introduction), [httpx](https://www.python-httpx.org/), or different LLM clients. Validation is performed with [Pydantic](https://docs.pydantic.dev/latest/).\n\n\u003cimg src=\"assets/architecture.jpg\" alt=\"Architecture Diagram\" style=\"width: 700px;\"\u003e\n\n- Language Agents natively support OpenAI, Anthropic, Gemini, Ollama, vLLM, Gradio, etc\n- Motor Agents natively support OpenVLA, RT1(upcoming)\n- Sensory Agents support Depth Anything, YOLO, Segment Anything 2\n\nJump to [getting started](#getting-started) to get up and running on [real hardware](https://colab.research.google.com/drive/1KN0JohcjHX42wABBHe-CxXP-NWJjbZts?usp=sharing) or [simulation](https://colab.research.google.com/drive/1gJlfEvsODZWGn3rK8Nx4A0kLnLzJtJG_?usp=sharing). Be sure to join our [Discord](https://discord.gg/BPQ7FEGxNb) for 🥇-winning discussions :)\n\n**⭐ Give us a star on GitHub if you like us!**\n\n### Motivation\n\n\u003cdetails\u003e\n\u003csummary\u003eThere is a significant barrier to entry for running SOTA models in robotics\u003c/summary\u003e\n\nIt is currently unrealistic to run state-of-the-art AI models on edge devices for responsive, real-time applications. Furthermore,\nthe complexity of integrating multiple models across different modalities is a significant barrier to entry for many researchers,\nhobbyists, and developers. This library aims to address these challenges by providing a simple, extensible, and efficient way to\nintegrate large models into existing robot stacks.\n\n\u003c/details\u003e\n\n### Goals\n\n\u003cdetails\u003e\n\u003csummary\u003eFacilitate data-collection and sharing among roboticists.\u003c/summary\u003e\n\nThis requires reducing much of the complexities involved with setting up inference endpoints, converting between different model formats, and collecting and storing new datasets for future availability.\n\nWe aim to achieve this by:\n\n1. Providing simple, Python-first abstrations that are modular, extensible and applicable to a wide range of tasks.\n2. Providing endpoints, weights, and interactive Gradio playgrounds for easy access to state-of-the-art models.\n3. Ensuring that this library is observation and action-space agnostic, allowing it to be used with any robot stack.\n\nBeyond just improved robustness and consistency, this architecture makes asynchronous and remote agent execution exceedingly simple. In particular we demonstrate how responsive natural language interactions can be achieved in under 30 lines of Python code.\n\n\u003c/details\u003e\n\n### Limitations\n\n_Embodied Agents are not yet capable of learning from in-context experience_:\n\n- Frameworks for advanced RAG techniques are clumsy at best for OOD embodied applications however that may improve.\n- Amount of data required for fine-tuning is still prohibitively large and expensive to collect.\n- Online RL is still in its infancy and not yet practical for most applications.\n\n### Scope\n\n- This library is intended to be used for research and prototyping.\n- This library is still experimental and under active development. Breaking changes may occur although they will be avoided as much as possible. Feel free to report issues!\n\n### Features\n\n- Extensible, user-friendly Python SDK with explicit typing and modularity\n- Asynchronous and remote thread-safe agent execution for maximal responsiveness and scalability.\n- Full-compatibility with HuggingFace Spaces, Datasets, Gymnasium Spaces, Ollama, and any OpenAI-compatible API.\n- Automatic dataset-recording and optionally uploads dataset to huggingface hub.\n\n### Endpoints\n\n- [OpenVLA](https://api.mbodi.ai/community-models/)\n- [Sensory Tools](https://api.mbodi.ai/sense/) (Depth Estimation, Image Segmentation, Object Detection)\n\n### Roadmap\n\n- [x] OpenVLA Motor Agent\n- [x] Automatic dataset recording on Robot\n- [x] Yolo, SAM2, DepthAnything Sensory Agents\n- [x] Auto Agent\n- [x] Google Gemini Backend\n- [x] Retrieval-Augmented Generation (RAG)\n- [ ] Pi0 Motor Agent\n- [ ] ROS integration\n- [ ] More Motor Agents, e.g., RT1, Octo\n- [ ] More device support, e.g., OpenCV camera\n- [ ] Fine-tuning Scripts\n\n## Installation\n\n```shell\npip install mbodied\n\n# With extra dependencies, e.g., torch, opencv-python, etc.\npip install mbodied[extras]\n\n# For audio support\npip install mbodied[audio]\n```\n\nOr install from source:\n\n```shell\npip install git+https://github.com/mbodiai/embodied-agents.git\n```\n\n## Getting Started\n\n### Customize a Motion to fit a robot's action space.\n\n```python\nfrom mbodied.types.motion.control import HandControl, FullJointControl\nfrom mbodied.types.motion import AbsoluteMotionField, RelativeMotionField\n\nclass FineGrainedHandControl(HandControl):\n    comment: str = Field(None, description=\"A comment to voice aloud.\")\n    index: FullJointControl = AbsoluteMotionField([0,0,0],bounds=[-3.14, 3.14], shape=(3,))\n    thumb: FullJointControl = RelativeMotionField([0,0,0],bounds=[-3.14, 3.14], shape=(3,))\n```\n\n### Run a robotics transformer model on a robot.\n\n```python\nimport os\nfrom mbodied.agents import LanguageAgent\nfrom mbodied.agents.motion import OpenVlaAgent\nfrom mbodied.agents.sense.audio import AudioAgent\nfrom mbodied.robots import SimRobot\n\ncognition = LanguageAgent(\n  context=\"You are an embodied planner that responds with a python list of strings and nothing else.\",\n  api_key=os.getenv(\"OPENAI_API_KEY\"),\n  model_src=\"openai\",\n  recorder=\"auto\",\n)\naudio = AudioAgent(use_pyaudio=False, api_key=os.getenv(\"OPENAI_API_KEY\")) # pyaudio is buggy on mac\nmotion = OpenVlaAgent(model_src=\"https://api.mbodi.ai/community-models/\")\n\n# Subclass and override do() and capture() methods.\nrobot = SimRobot()\n\ninstruction = audio.listen()\nplan = cognition.act(instruction, robot.capture())\n\nfor step in plan.strip('[]').strip().split(','):\n  print(\"\\nMotor agent is executing step: \", step, \"\\n\")\n  for _ in range(10):\n    hand_control = motion.act(step, robot.capture())\n    robot.do(hand_control)\n```\n\nExample Scripts:\n\n- [1_simple_robot_agent.py](examples/1_simple_robot_agent.py): A very simple language based cognitive agent taking instruction from user and output voice and actions.\n- [2_openvla_motor_agent_example.py](examples/2_openvla_motor_agent_example.py): Run robotic transformers, i.e. OpenVLA, in several lines on the robot.\n- [3_reason_plan_act_robot.py](examples/3_reason_plan_act_robot.py): Full example of language based cognitive agent and OpenVLA motor agent executing task.\n- [4_language_reason_plan_act_robot.py](examples/4_language_reason_plan_act_robot.py): Full example of all languaged based cognitive and motor agent executing task.\n- [5_teach_robot_record_dataset.py](examples/5_teach_robot_record_dataset.py): Example of collecting dataset on robot's action at a specific frequency by just yelling at the robot!\n- [6_robot_with_rag.py](examples/6_robot_with_rag.py): Example with RAG to retrieve skills to run on the robot.\n\n### Notebooks\n\nReal Robot Hardware: [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1KN0JohcjHX42wABBHe-CxXP-NWJjbZts?usp=sharing)\n\nSimulation with: [SimplerEnv](https://github.com/simpler-env/SimplerEnv.git) : [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/18oiuw1yTxO5x-eT7Z8qNyWtjyYd8cECI?usp=sharing)\n\nRun OpenVLA with embodied-agents in simulation: [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1flnMrqyepGOO8J9rE6rehzaLdZPsw6lX?usp=sharing)\n\nRecord dataset on a robot: [\u003cimg align=\"center\" src=\"https://colab.research.google.com/assets/colab-badge.svg\" /\u003e](https://colab.research.google.com/drive/15UuFbMUJGEjqJ_7I_b5EvKvLCKnAc8bB?usp=sharing)\n\n## The [Sample](mbodied/base/sample.py) Class\n\nThe Sample class is a base model for serializing, recording, and manipulating arbitrary data. It is designed to be extendable, flexible, and strongly typed. By wrapping your observation or action objects in the [Sample](mbodied/base/sample.py) class, you'll be able to convert to and from the following with ease:\n\n- A Gym space for creating a new Gym environment.\n- A flattened list, array, or tensor for plugging into an ML model.\n- A HuggingFace dataset with semantic search capabilities.\n- A Pydantic BaseModel for reliable and quick json serialization/deserialization.\n\nTo learn more about all of the possibilities with embodied agents, check out the [documentation](https://mbodi-ai-mbodied-agents.readthedocs-hosted.com/en/latest/)\n\n### 💡 Did you know\n\n- You can `pack` a list of `Sample`s or Dicts into a single `Sample` or `Dict` and `unpack` accordingly?\n- You can `unflatten` any python structure into a `Sample` class so long you provide it with a valid json schema?\n\n## API Reference\n\n#### Creating a Sample\n\nCreating a Sample requires just wrapping a python dictionary with the `Sample` class. Additionally, they can be made from kwargs, Gym Spaces, and Tensors to name a few.\n\n```python\nfrom mbodied.types.sample import Sample\n# Creating a Sample instance\nsample = Sample(observation=[1,2,3], action=[4,5,6])\n\n# Flattening the Sample instance\nflat_list = sample.flatten()\nprint(flat_list) # Output: [1, 2, 3, 4, 5, 6]\n\n# Generating a simplified JSON schema\n\u003e\u003e\u003e schema = sample.schema()\n{'type': 'object', 'properties': {'observation': {'type': 'array', 'items': {'type': 'integer'}}, 'action': {'type': 'array', 'items': {'type': 'integer'}}}}\n\n# Unflattening a list into a Sample instance\nSample.unflatten(flat_list, schema)\n\u003e\u003e\u003e Sample(observation=[1, 2, 3], action=[4, 5, 6])\n```\n\n#### Serialization and Deserialization with Pydantic\n\nThe Sample class leverages Pydantic's powerful features for serialization and deserialization, allowing you to easily convert between Sample instances and JSON.\n\n```python\n# Serialize the Sample instance to JSON\nsample = Sample(observation=[1,2,3], action=[4,5,6])\njson_data = sample.model_dump_json()\nprint(json_data) # Output: '{\"observation\": [1, 2, 3], \"action\": [4, 5, 6]}'\n\n# Deserialize the JSON data back into a Sample instance\njson_data = '{\"observation\": [1, 2, 3], \"action\": [4, 5, 6]}'\nsample = Sample.model_validate(from_json(json_data))\nprint(sample) # Output: Sample(observation=[1, 2, 3], action=[4, 5, 6])\n```\n\n#### Converting to Different Containers\n\n```python\n# Converting to a dictionary\nsample_dict = sample.to(\"dict\")\nprint(sample_dict) # Output: {'observation': [1, 2, 3], 'action': [4, 5, 6]}\n\n# Converting to a NumPy array\nsample_np = sample.to(\"np\")\nprint(sample_np) # Output: array([1, 2, 3, 4, 5, 6])\n\n# Converting to a PyTorch tensor\nsample_pt = sample.to(\"pt\")\nprint(sample_pt) # Output: tensor([1, 2, 3, 4, 5, 6])\n```\n\n#### Gym Space Integration\n\n```python\ngym_space = sample.space()\nprint(gym_space)\n# Output: Dict('action': Box(-inf, inf, (3,), float64), 'observation': Box(-inf, inf, (3,), float64))\n```\n\nSee [sample.py](mbodied/base/sample.py) for more details.\n\n### Message\n\nThe [Message](mbodied/types/message.py) class represents a single completion sample space. It can be text, image, a list of text/images, Sample, or other modality. The Message class is designed to handle various types of content and supports different roles such as user, assistant, or system.\n\nYou can create a `Message` in versatile ways. They can all be understood by mbodi's backend.\n\n```python\nfrom mbodied.types.message import Message\n\nMessage(role=\"user\", content=\"example text\")\nMessage(role=\"user\", content=[\"example text\", Image(\"example.jpg\"), Image(\"example2.jpg\")])\nMessage(role=\"user\", content=[Sample(\"Hello\")])\n```\n\n### Backend\n\nThe [Backend](mbodied/base/backend.py) class is an abstract base class for Backend implementations. It provides the basic structure and methods required for interacting with different backend services, such as API calls for generating completions based on given messages. See [backend directory](mbodied/agents/backends) on how various backends are implemented.\n\n### Agent\n\n[Agent](mbodied/base/agent.py) is the base class for various agents listed below. It provides a template for creating agents that can talk to a remote backend/server and optionally record their actions and observations.\n\n### Language Agent\n\nThe [Language Agent](mbodied/agents/language/language_agent.py) can connect to different backends or transformers of your choice. It includes methods for recording conversations, managing context, looking up messages, forgetting messages, storing context, and acting based on an instruction and an image.\n\nNatively supports API services: OpenAI, Anthropic, Gemini vLLM, Ollama, HTTPX, or any gradio endpoints. More upcoming!\n\nTo use OpenAI for your robot backend:\n\n```python\nfrom mbodied.agents.language import LanguageAgent\n\nagent = LanguageAgent(context=\"You are a robot agent.\", model_src=\"openai\")\n```\n\nTo execute an instruction:\n\n```python\ninstruction = \"pick up the fork\"\nresponse = robot_agent.act(instruction, image)\n```\n\nLanguage Agent can connect to vLLM as well. For example, suppose you are running a vLLM server Mistral-7B on 1.2.3.4:1234. All you need to do is:\n\n```python\nagent = LanguageAgent(\n    context=context,\n    model_src=\"openai\",\n    model_kwargs={\"api_key\": \"EMPTY\", \"base_url\": \"http://1.2.3.4:1234/v1\"},\n)\nresponse = agent.act(\"Hello, how are you?\", model=\"mistralai/Mistral-7B-Instruct-v0.3\")\n```\n\nExample using Ollama:\n\n```python\nagent = LanguageAgent(\n    context=\"You are a robot agent.\", model_src=\"ollama\",\n    model_kwargs={\"endpoint\": \"http://localhost:11434/api/chat\"}\n)\nresponse = agent.act(\"Hello, how are you?\", model=\"llama3.1\")\n```\n\n### Motor Agent\n\n[Motor Agent](mbodied/agents/motion/motor_agent.py) is similar to Language Agent but instead of returning a string, it always returns a `Motion`. Motor Agent is generally powered by robotic transformer models, e.g., OpenVLA, RT1, Octo, etc.\nSome small models, like RT1, can run on edge devices. However, some, like OpenVLA, may be challenging to run without quantization. See [OpenVLA Agent](mbodied/agents/motion/openvla_agent.py) and an [example OpenVLA server](examples/servers/gradio_example_openvla.py)\n\n### Sensory Agent\n\nThese agents interact with the environment to collect sensor data. They always return a `SensorReading`, which can be various forms of processed sensory input such as images, depth data, or audio signals.\n\nCurrently, we have:\n\n- [depth estimation](mbodied/agents/sense/depth_estimation_agent.py)\n- [object detection](mbodied/agents/sense/object_detection_agent.py)\n- [image segmentation](mbodied/agents/sense/segmentation_agent.py)\n\nagents that process robot's sensor information.\n\n### Auto Agent\n\n[Auto Agent](mbodied/agents/auto/auto_agent.py) dynamically selects and initializes the correct agent based on the task and model.\n\n```python\nfrom mbodied.agents.auto.auto_agent import AutoAgent\n\n# This makes it a LanguageAgent\nagent = AutoAgent(task=\"language\", model_src=\"openai\")\nresponse = agent.act(\"What is the capital of France?\")\n\n# This makes it a motor agent: OpenVlaAgent\nauto_agent = AutoAgent(task=\"motion-openvla\", model_src=\"https://api.mbodi.ai/community-models/\")\naction = auto_agent.act(\"move hand forward\", Image(size=(224, 224)))\n\n# This makes it a sensory agent: DepthEstimationAgent\nauto_agent = AutoAgent(task=\"sense-depth-estimation\", model_src=\"https://api.mbodi.ai/sense/\")\ndepth = auto_agent.act(image=Image(size=(224, 224)))\n```\n\nAlternatively, you can use `get_agent` method in [auto_agent](mbodied/agents/auto/auto_agent.py) as well.\n\n```python\nlanguage_agent = get_agent(task=\"language\", model_src=\"openai\")\n```\n\n### Motions\n\nThe [motion_controls](mbodied/types/motion_controls.py) module defines various motions to control a robot as Pydantic models. They are also subclassed from `Sample`, thus possessing all the capability of `Sample` as mentioned above. These controls cover a range of actions, from simple joint movements to complex poses and full robot control.\n\n### Robot\n\nYou can integrate your custom robot hardware by subclassing [Robot](mbodied/robot/robot.py) quite easily. You only need to implement `do()` function to perform actions (and some additional methods if you want to record dataset on the robot). In our examples, we use a [mock robot](mbodied/robot/sim_robot.py). We also have an [XArm robot](mbodied/robot/xarm_robot.py) as an example.\n\n#### Recording a Dataset\n\nRecording a dataset on a robot is very easy! All you need to do is implement the `get_observation()`, `get_state()`, and `prepare_action()` methods for your robot. After that, you can record a dataset on your robot anytime you want. See [examples/5_teach_robot_record_dataset.py](examples/5_teach_robot_record_dataset.py) and this colab: [\u003cimg align=\"center\" src=\"https://colab.research.google.com/assets/colab-badge.svg\" /\u003e](https://colab.research.google.com/drive/15UuFbMUJGEjqJ_7I_b5EvKvLCKnAc8bB?usp=sharing) for more details.\n\n```python\nfrom mbodied.robots import SimRobot\nfrom mbodied.types.motion.control import HandControl, Pose\n\nrobot = SimRobot()\nrobot.init_recorder(frequency_hz=5)\nwith robot.record(\"pick up the fork\"):\n  motion = HandControl(pose=Pose(x=0.1, y=0.2, z=0.3, roll=0.1, pitch=0.2, yaw=0.3))\n  robot.do(motion)\n```\n\n### Recorder\n\nDataset [Recorder](mbodied/data/recording.py) is a lower level recorder to record your conversation and the robot's actions to a dataset as you interact with/teach the robot. You can define any observation space and action space for the Recorder. See [gymnasium](https://github.com/Farama-Foundation/Gymnasium) for more details about spaces.\n\n```python\nfrom mbodied.data.recording import Recorder\nfrom mbodied.types.motion.control import HandControl\nfrom mbodied.types.sense.vision import Image\nfrom gymnasium import spaces\n\nobservation_space = spaces.Dict({\n    'image': Image(size=(224, 224)).space(),\n    'instruction': spaces.Text(1000)\n})\naction_space = HandControl().space()\nrecorder = Recorder('example_recorder', out_dir='saved_datasets', observation_space=observation_space, action_space=action_space)\n\n# Every time robot makes a conversation or performs an action:\nrecorder.record(observation={'image': image, 'instruction': instruction,}, action=hand_control)\n```\n\nThe dataset is saved to `./saved_datasets`.\n\n### Replayer\n\nThe [Replayer](mbodied/data/replaying.py) class is designed to process and manage data stored in HDF5 files generated by `Recorder`. It provides a variety of functionalities, including reading samples, generating statistics, extracting unique items, and converting datasets for use with HuggingFace. The Replayer also supports saving specific images during processing and offers a command-line interface for various operations.\n\nExample for iterating through a dataset from Recorder with Replayer:\n\n```python\nfrom mbodied.data.replaying import Replayer\n\nreplayer = Replayer(path=str(\"path/to/dataset.h5\"))\nfor observation, action in replayer:\n   ...\n```\n\n## Directory Structure\n\n```shell\n├─ assets/ ............. Images, icons, and other static assets\n├─ examples/ ........... Example scripts and usage demonstrations\n├─ resources/ .......... Additional resources for examples\n├─ src/\n│  └─ mbodied/\n│     ├─ agents/ ....... Modules for robot agents\n│     │  ├─ backends/ .. Backend implementations for different services for agents\n│     │  ├─ language/ .. Language based agents modules\n│     │  ├─ motion/ .... Motion based agents modules\n│     │  └─ sense/ ..... Sensory, e.g. audio, processing modules\n│     ├─ data/ ......... Data handling and processing\n│     ├─ hardware/ ..... Hardware modules, i.e. camera\n│     ├─ robot/ ........ Robot interface and interaction\n│     └─ types/ ........ Common types and definitions\n└─ tests/ .............. Unit tests\n```\n\n## Contributing\n\nWe welcome issues, questions and PRs. See the [contributing guide](CONTRIBUTING.md) for more information.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmbodiai%2Fembodied-agents","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmbodiai%2Fembodied-agents","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmbodiai%2Fembodied-agents/lists"}