{"id":13826165,"url":"https://github.com/jaketae/storyteller","last_synced_at":"2025-04-05T00:10:12.390Z","repository":{"id":65673937,"uuid":"583616027","full_name":"jaketae/storyteller","owner":"jaketae","description":"Multimodal AI Story Teller, built with Stable Diffusion, GPT, and neural text-to-speech","archived":false,"fork":false,"pushed_at":"2023-08-29T16:31:11.000Z","size":4816,"stargazers_count":518,"open_issues_count":0,"forks_count":63,"subscribers_count":13,"default_branch":"master","last_synced_at":"2025-03-28T23:08:58.011Z","etag":null,"topics":["ddpm","diffusion-models","gpt","image-generation","natural-language-generation","pytorch","stable-diffusion","text-to-image","text-to-speech","text-to-video","video-generation"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/jaketae.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2022-12-30T10:31:29.000Z","updated_at":"2025-03-24T13:04:35.000Z","dependencies_parsed_at":"2024-01-15T16:21:57.310Z","dependency_job_id":"a5f5feb0-d1fb-4d39-a930-dde8a64ec6f2","html_url":"https://github.com/jaketae/storyteller","commit_stats":null,"previous_names":[],"tags_count":2,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jaketae%2Fstoryteller","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jaketae%2Fstoryteller/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jaketae%2Fstoryteller/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jaketae%2Fstoryteller/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/jaketae","download_url":"https://codeload.github.com/jaketae/storyteller/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247266565,"owners_count":20910836,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ddpm","diffusion-models","gpt","image-generation","natural-language-generation","pytorch","stable-diffusion","text-to-image","text-to-speech","text-to-video","video-generation"],"created_at":"2024-08-04T09:01:33.212Z","updated_at":"2025-04-05T00:10:12.370Z","avatar_url":"https://github.com/jaketae.png","language":"Python","funding_links":[],"categories":["Python","🤖 ChatGPT Agents"],"sub_categories":["Creative \u0026 Content"],"readme":"# StoryTeller\n\n[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/17C284MOUDQMxV6bRbgVRH4GXsb87iADW?usp=sharing)\n[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)\n[![pre-commit](https://img.shields.io/badge/pre--commit-enabled-green?logo=pre-commit\u0026logoColor=white)](https://github.com/pre-commit/pre-commit)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n\nA multimodal AI storyteller, built with [Stable Diffusion](https://huggingface.co/spaces/stabilityai/stable-diffusion), GPT, and neural text-to-speech (TTS).\n\nGiven a prompt as an opening line of a story, GPT writes the rest of the plot; Stable Diffusion draws an image for each sentence; a TTS model narrates each line, resulting in a fully animated video of a short story, replete with audio and visuals.\n\n\u003cimg id=\"default-output\" src=\"https://user-images.githubusercontent.com/25360440/210071764-51ed5872-ba56-4ed0-919b-d9ce65110185.gif\" alt=\"Example output generated with the default prompt.\"\u003e\n\n## Installation\n\n### PyPI\n\nStory Teller is available on [PyPI](https://pypi.org/project/storyteller-core/).\n\n```\n$ pip install storyteller-core\n```\n\n### Source\n\n1. Clone the repository.\n\n```\n$ git clone https://github.com/jaketae/storyteller.git\n$ cd storyteller\n```\n\n2. Install dependencies.\n\n```\n$ pip install .\n```\n\n\u003e [!NOTE]\n\u003e For Apple Silicon users, [`mecab-python3`](https://github.com/SamuraiT/mecab-python3) is not available. You need to install `mecab` before running `pip install`. You can do this with [Hombrew](https://www.google.com/search?client=safari\u0026rls=en\u0026q=homebrew\u0026ie=UTF-8\u0026oe=UTF-8) via `brew install mecab`. For more information, refer to https://github.com/SamuraiT/mecab-python3/issues/84.\n\n3. (Optional) To develop locally, install `dev` dependencies and install pre-commit hooks. This will automatically trigger linting and code quality checks before each commit.\n\n```\n$ pip install -e .[dev]\n$ pre-commit install\n```\n\n## Quickstart\n\nThe quickest way to run a demo is by using the command line interface (CLI). To get started, simply type:\n\n```\n$ storyteller\n```\n\nThis command will initialize the story with the default prompt of `Once upon a time, unicorns roamed the Earth`. An\nexample of the output that will be generated [can be seen in the animation above](#default-output).\nYou can customize the beginning of your story by using the `--writer_prompt` argument. For example, if you would like to\nstart your story with the text `The ravenous cat, driven by an insatiable craving for tuna, devised a daring plan to break into the local fish market's coveted tuna reserve.`,\nyour CLI command would look as follows:\n\n```\nstoryteller --writer_prompt \"The ravenous cat, driven by an insatiable craving for tuna, devised a daring plan to break into the local fish market's coveted tuna reserve.\"\n```\n\nThe final video will be saved in the `/out/out.mp4` directory, along with other intermediate files such as images,\naudio files, and subtitles.\n\nTo adjust the default settings with custom parameters, you can use the different CLI flags as needed. To see a list of\nall available options, type:\n\n```\n$ storyteller --help\n```\n\nThis will provide you with a list of the options, their descriptions and their defaults.\n\n\n```\noptions:\n  -h, --help            show this help message and exit\n  --writer_prompt WRITER_PROMPT\n                        The prompt to be used for the writer model. This is the text with which your story will begin. Default:\n                        'Once upon a time, unicorns roamed the Earth.'\n  --painter_prompt_prefix PAINTER_PROMPT_PREFIX\n                        The prefix to be used for the painter model's prompt. Default: 'Beautiful painting'\n  --num_images NUM_IMAGES\n                        The number of images to be generated. Those images will be composed in sequence into a video. Default:\n                        10\n  --output_dir OUTPUT_DIR\n                        The directory to save the generated files to. Default: 'out'\n  --seed SEED           The seed value to be used for randomization. Default: 42\n  --max_new_tokens MAX_NEW_TOKENS\n                        Maximum number of new tokens to generate in the writer model. Default: 50\n  --writer WRITER       Text generation model to use. Default: 'gpt2'\n  --painter PAINTER     Image generation model to use. Default: 'stabilityai/stable-diffusion-2'\n  --speaker SPEAKER     Text-to-speech (TTS) generation model. Default: 'tts_models/en/ljspeech/glow-tts'\n  --writer_device WRITER_DEVICE\n                        Text generation device to use. Default: 'cpu'\n  --painter_device PAINTER_DEVICE\n                        Image generation device to use. Default: 'cpu'\n  --writer_dtype WRITER_DTYPE\n                        Text generation dtype to use. Default: 'float32'\n  --painter_dtype PAINTER_DTYPE\n                        Image generation dtype to use. Default: 'float32'\n  --enable_attention_slicing ENABLE_ATTENTION_SLICING\n                        Whether to enable attention slicing for diffusion. Default: 'False'\n```\n\n## Usage\n\n### Command Line Interface\n\n#### CUDA\n\nIf you have a CUDA-enabled machine, run\n\n```\n$ storyteller --writer_device cuda --painter_device cuda\n```\n\nto utilize GPU.\n\nYou can also place each model on separate devices if loading all models on a single device exceeds available VRAM.\n\n```\n$ storyteller --writer_device cuda:0 --painter_device cuda:1\n```\n\n$ For faster generation, consider using half-precision.\n\n```\n$ storyteller --writer_device cuda --painter_device cuda --writer_dtype float16 --painter_dtype float16\n```\n\n#### Apple Silicon\n\n\u003e [!NOTE]\n\u003e PyTorch support for Apple Silicon ([MPS](https://pytorch.org/docs/stable/notes/mps.html)) is work in progress. At the time of writing, `torch.cumsum` does not work with `torch.int64` ([issue](https://github.com/pytorch/pytorch/issues/96610)) on PyTorch stable 2.0.1; it works on nightly only.\n\nIf you are on an Apple Silicon machine, run\n\n```\n$ storyteller --writer_device mps --painter_device mps\n```\n\nif you want to use MPS acceleration for both models.\n\nFor faster generation, consider enabling [attention-slicing](https://huggingface.co/docs/diffusers/optimization/fp16#sliced-attention-for-additional-memory-savings) to save on memory.\n\n```\n$ storyteller --enable_attention_slicing true\n```\n\n### Python\n\nFor more advanced use cases, you can also directly interface with Story Teller in Python code.\n\n1. Load the model with defaults.\n\n```python\nfrom storyteller import StoryTeller\n\nstory_teller = StoryTeller.from_default()\nstory_teller.generate(...)\n```\n\n2. Alternatively, configure the model with custom settings.\n\n```python\nfrom storyteller import StoryTeller, StoryTellerConfig\n\nconfig = StoryTellerConfig(\n    writer=\"gpt2-large\",\n    painter=\"CompVis/stable-diffusion-v1-4\",\n    max_new_tokens=100,\n)\n\nstory_teller = StoryTeller(config)\nstory_teller.generate(...)\n```\n\n## License\n\nReleased under the [MIT License](LICENSE).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjaketae%2Fstoryteller","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjaketae%2Fstoryteller","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjaketae%2Fstoryteller/lists"}