{"id":51717089,"url":"https://github.com/zanni098/ai-documentary-101","last_synced_at":"2026-07-17T04:37:19.182Z","repository":{"id":369900704,"uuid":"1292144686","full_name":"zanni098/ai-documentary-101","owner":"zanni098","description":"Full reproducible source for the 'AI Documentary Making 101' short film: prompts, scripts, and assets to recreate it end to end.","archived":false,"fork":false,"pushed_at":"2026-07-07T10:34:23.000Z","size":17442,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"master","last_synced_at":"2026-07-07T12:16:02.477Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/zanni098.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-07-07T10:19:09.000Z","updated_at":"2026-07-07T10:34:32.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/zanni098/ai-documentary-101","commit_stats":null,"previous_names":["zanni098/ai-documentary-101"],"tags_count":1,"template":false,"template_full_name":null,"purl":"pkg:github/zanni098/ai-documentary-101","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zanni098%2Fai-documentary-101","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zanni098%2Fai-documentary-101/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zanni098%2Fai-documentary-101/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zanni098%2Fai-documentary-101/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/zanni098","download_url":"https://codeload.github.com/zanni098/ai-documentary-101/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zanni098%2Fai-documentary-101/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35568173,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-17T02:00:06.162Z","response_time":116,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2026-07-17T04:37:18.510Z","updated_at":"2026-07-17T04:37:19.172Z","avatar_url":"https://github.com/zanni098.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# AI Documentary Making 101 — Full Project Source\n\n**A ~6-minute documentary — script, narrator, exhibits, voice, SFX, and edit — made entirely with AI tools, and directed almost entirely by Claude Code.** This repo is the complete, reproducible source: every prompt, every script, every asset, in the exact order they were used to build the final video.\n\nFinal video: **[youtube.com/watch?v=uPMaQBqYsCw](https://www.youtube.com/watch?v=uPMaQBqYsCw)**\n\n---\n\n## What's in this repo\n\n```\ndocs/                       planning + prompts (the \"how\")\n  DOCUMENTARY_PLAN.md          original production plan (concept, tone, early draft script)\n  SCRIPT_v2_desert.md          the SHOOTING SCRIPT actually used (desert/helicopter structure)\n  CHARACTER_SHOTS_veo.md       the narrator shot list + non-repetitive cutting map\n  VOX_EXHIBIT_PROMPTS.md       verbatim ChatGPT prompts for all 9 Vox exhibit clips\n  VO_SCRIPT_SEGMENTS.md        original voiceover segment map (SEG_01-21)\n  GEMINI_VIDEO_PROMPTS.md      copy-paste Gemini/Veo prompts (fallback path, pre-Vertex AI)\nscripts/                     the actual assembly pipeline (run in this order)\n  build_film.py                 renders all 32 narrative beats (VO + character/exhibit clips + SFX)\n  fix_sync_beats.py             re-renders beats where VO was shorter than the clip (removes sped-up motion)\n  fix_clip_swap.py              swaps in a more natural walking/gesture clip for 5 beats\npreview/\n  preview_720p.mp4              compressed 720p preview of the final cut (full quality is on YouTube)\n```\n\nBinary assets (character images, Veo clips, Vox exhibit videos, voice, SFX, music, and the pre-rendered beats) are **not** committed to git — they're attached as [GitHub Release](../../releases) downloads instead, so the repo itself stays small and fast to clone. See **Asset bundles** below.\n\n---\n\n## Prerequisites\n\n| Need | Used for | Tier |\n|---|---|---|\n| Google Cloud project + Vertex AI enabled | Veo 3 (character clips) | Billed, but Veo is the only paid step here |\n| ElevenLabs account | Voice (narration) + sound-effects generation | Free tier is enough for ~20 short lines |\n| ChatGPT (GPT Image) | Vox exhibit stills | Free tier, ~1-2 images per ~30 min |\n| Higgsfield.ai account | Vox exhibit animation (image→video) | Free — use the **Unlimited toggle** models only (\"Enhanced Seedance 2.0 Fast\"), never paid credits |\n| ffmpeg + ffprobe | All video/audio assembly | Free |\n| Python 3 | Running the assembly scripts | Free |\n| CapCut (desktop) | Final timeline assembly + export | Free |\n\n---\n\n## Step 1 — Character design\n\nThe masked narrator (\"The Boring Studio\") started from two user-supplied reference images (included in the asset bundle as `character/`):\n- A full character turnaround/design sheet (palette `#0D0D0D` / `#1A1A1A` / `#2B2B2B` / `#E04A6A` / `#FF7A00` / rust, materials called out on the sheet itself)\n- A single canonical reference photo (`Screenshot (498).png`) used as the exact image-to-video seed for every Veo shot, to guarantee the mask/suit/tie design never drifts between clips\n\nIf regenerating this character from scratch (no reference images), the working verbal description used throughout is:\n\u003e *\"A man in a tailored matte-black suit, black leather gloves, a rusted orange-brown metal mask fully hiding his face with narrow dark eye-slits (mesh, not glass), black shirt, dark patterned tie.\"*\n\n## Step 2 — Generate the narrator (character) clips\n\nAll narrator shots are **image-to-video** (not text-to-video) from the single reference photo above — this is the trick that keeps the mask/suit design identical across every clip. Generated via **Vertex AI Veo 3**, model `veo-3.0-fast-generate-001`:\n\n```\nPOST https://{LOCATION}-aiplatform.googleapis.com/v1/projects/{PROJECT}/locations/{LOCATION}/publishers/google/models/veo-3.0-fast-generate-001:predictLongRunning\n{\n  \"instances\": [{\n    \"prompt\": \"\u003cshot description, see below\u003e\",\n    \"image\": {\"bytesBase64Encoded\": \"\u003cbase64 of Screenshot (498).png\u003e\", \"mimeType\": \"image/png\"}\n  }]\n}\n```\nThen poll the same model's `:fetchPredictOperation` until done; the result video is at `response.videos[0].bytesBase64Encoded`.\n\nShot list (full descriptions + the non-repetitive cutting map are in `docs/CHARACTER_SHOTS_veo.md`):\n\n| File | Shot |\n|---|---|\n| `A1_aerial` | Aerial over dunes, helicopter banking low (reused as-is throughout) |\n| `A2_heli_landing` | Helicopter lands, dust rolls toward camera |\n| `A3_steps_out_v3` | Hero shot — steps out of the helicopter, walks forward, stops |\n| `R1_closeup_v3` | Extreme close-up on the mask, slow push-in |\n| `R2_front_medium_v3` | Waist-up \"main talking\" shot, gesturing |\n| `R3_wide_power_v3` | Wide shot — helicopter, guards, SUVs visible behind him |\n| `R4_turn_v3` | 3/4 turn shot *(superseded — see Step 2b)* |\n| `R5_low_angle_v3` | Low angle, powerful/still |\n| `R7_walk_v3` | Slow walk toward camera |\n| `R8_golden_v3` | Golden-hour wide, long shadows (reflection scene) |\n| `A9_departure_v3` | Walks back to the helicopter, lifts off into the sunset |\n\n## Step 2b — The natural-motion fix\n\n`R4_turn_v3` looked stiff/robotic on review. It was replaced everywhere it was used with a hand-picked 7-second segment (walking + natural hand/head gesture, skipping the helicopter-descent portion) extracted from a separately user-generated reference clip:\n\n```bash\nffmpeg -i \"source_clip.mp4\" -ss 3.0 -t 7.0 -c:v libx264 -crf 16 -an RNEW_walktalk_v1.mp4\n```\nThis `RNEW_walktalk_v1.mp4` is included in the asset bundle. `scripts/fix_clip_swap.py` re-renders the 5 affected beats using it instead of `R4_turn_v3`.\n\n## Step 3 — Generate the 9 Vox exhibit clips\n\nThe \"evidence\" clips (Vox-style infographic explainer, warm paper/collage look) are a **separate visual world** from the narrator — full prompts (verbatim ChatGPT conversations for both the still image and the Higgsfield animation prompt) are in `docs/VOX_EXHIBIT_PROMPTS.md`.\n\nPipeline per exhibit: ChatGPT (GPT Image) generates the still → the still is fed to **Higgsfield \"Enhanced Seedance 2.0 Fast\"** (image-to-video, **Unlimited toggle only — never paid credits**) with a detailed shot-by-shot animation prompt → the resulting 15s 720p clip is upscaled to 2560×1440 (`master/vox_2k/VOX_01...VOX_09`).\n\nTopics, in order: future of filmmaking → old workflow vs AI → Remotion+Higgsfield overview → Remotion deeper → Remotion's limits → Higgsfield deeper → Higgsfield's library → Higgsfield's limits/cost → Claude Code orchestration (the reveal).\n\n## Step 4 — Voiceover\n\nElevenLabs, voice **Brian** (`nPczCjzI2devNBz1zQrb` — \"Deep, Resonant and Comforting\"), model `eleven_multilingual_v2`:\n```json\n{\"stability\": 0.40, \"similarity_boost\": 0.85, \"style\": 0.30, \"use_speaker_boost\": true, \"speed\": 0.90}\n```\nLine-by-line script is in `docs/SCRIPT_v2_desert.md` (current) — `docs/VO_SCRIPT_SEGMENTS.md` has the original numbering (`SEG_04`-`SEG_21` are reused as-is; `SEG_V2_01`-`03` are new desert-arrival lines that replaced the old cold open).\n\n## Step 5 — SFX\n\nElevenLabs sound-generation API. Cues used: rotor spin-up/spin-down, helicopter idle/takeoff, footsteps in sand, suit rustle, desert wind bed (looped, low), desert ambience, two sub-boom impacts (title card + reveal beat). All included as `sfx/*.mp3` in the asset bundle.\n\n## Step 6 — Music\n\n6 pre-composed 30-second cues (sub-bass drone → exposition pulse → rising reveal → tension drone → reflective piano → climax swell), included as `music/*.mp3`. These are licensed/owned tracks — swap in your own if reproducing this publicly.\n\n## Step 7 — Assemble\n\nRun in this exact order from inside the folder containing `master/`, `narrator/`:\n\n```bash\npython scripts/build_film.py        # renders all 32 beats (VO + clips + SFX) into master/final_scenes/\npython scripts/fix_sync_beats.py    # re-renders beats where VO was shorter than the clip, at natural speed\npython scripts/fix_clip_swap.py     # swaps R4_turn_v3 -\u003e RNEW_walktalk_v1 in the 5 beats that used it\n```\n\nThen fix the exhibit/narration order (two Vox clips play out of order relative to the narration otherwise — see below) and build a correctly-numbered copy for import:\n\n```python\nORDER = [\n    'S1a_aerial','S1b_title','S1c_landing','S1d_stepsout',\n    'S2a_vo1','S2b_vo2','S2c_vo3',\n    'S3a_vo04','S3d_vox2','S3c_vo05','S3b_vox1','S3e_vo06',      # vox1/vox2 swapped\n    'S4a_vo07','S4b_vox3','S4c_vo08','S4d_vox4','S4f_vo09','S4e_vox5','S4h_vo10','S4g_vox6','S4i_vox7',  # vox5/vo09 and vox6/vo10 swapped\n    'S5a_vo11','S5b_vox8',\n    'S6a_vo12','S6b_vo13','S6c_vox9','S6d_vo14',\n    'S7a_vo15','S7b_vo16','S7c_vo17',\n    'S8a_vo18','S8b_vo19','S8c_departure',\n]\n```\n*(Why: the narrator would say \"it has a ceiling\" or \"Higgsfield\" a beat **after** the exhibit that illustrates it had already played. The fix reorders so narration always introduces a topic before its exhibit shows it.)*\n\nA separate music timeline (6 cues, crossfaded, matched to the film's exact runtime) is built independently — voice+video+SFX and music are always kept as separate deliverables so music can be re-balanced without re-rendering anything.\n\n## Step 8 — Final cut (CapCut)\n\n1. New project → import all 33 ordered beat clips + the music track.\n2. Drag all 33 clips onto the timeline in numeric order (they concatenate automatically, no gaps).\n3. Music on its own track, same start point, volume **-15dB** (sits under dialogue/SFX, doesn't compete).\n4. Export 1080p, H.264, mp4.\n\n---\n\n## Asset bundles ([v1.0 release](https://github.com/zanni098/ai-documentary-101/releases/tag/v1.0))\n\n| Bundle | Contents | Size |\n|---|---|---|\n| [`core-assets.zip`](https://github.com/zanni098/ai-documentary-101/releases/download/v1.0/core-assets.zip) | Character images + all 12 narrator Veo clips + all 9 Vox exhibit clips (2K) + voice + SFX + music + the 33 pre-rendered final beats | ~590MB |\n| [`archive-exploration.zip`](https://github.com/zanni098/ai-documentary-101/releases/download/v1.0/archive-exploration.zip) | Early sample renders, the rejected Remotion-only attempt, superseded (wrong-character) narrator test clips | ~335MB |\n| [`archive-vox-source.zip`](https://github.com/zanni098/ai-documentary-101/releases/download/v1.0/archive-vox-source.zip) | Original 720p Vox stills + raw Higgsfield renders before 2K upscale (the source for `docs/VOX_EXHIBIT_PROMPTS.md`) | ~280MB |\n\nGrab `core-assets.zip` if you just want to reproduce the pipeline above. The archive bundles are kept for transparency/history only — not needed to rebuild the video.\n\n---\n\n## Credits\n\nBuilt with Claude Code (Anthropic), Google Vertex AI (Veo 3), ElevenLabs, ChatGPT (GPT Image), Higgsfield.ai, ffmpeg, and CapCut.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzanni098%2Fai-documentary-101","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fzanni098%2Fai-documentary-101","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzanni098%2Fai-documentary-101/lists"}