{"id":51617521,"url":"https://github.com/tech-sumit/liveavatarstream3d","last_synced_at":"2026-07-12T15:01:19.269Z","repository":{"id":369258418,"uuid":"1273699896","full_name":"tech-sumit/LiveAvatarStream3D","owner":"tech-sumit","description":"🎬 Browser-based 3D talking-avatar studio — write a script, a lip-synced 3D presenter delivers it, export MP4 client-side with WebCodecs. Direction as data. Agent-native via WebMCP. No render server.","archived":false,"fork":false,"pushed_at":"2026-07-04T11:13:26.000Z","size":19260,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-04T12:11:10.840Z","etag":null,"topics":["3d","ai-agents","avatar","cloudflare-workers","elevenlabs","lip-sync","mcp","talking-head","text-to-speech","threejs","typescript","video-generation","virtual-presenter","webcodecs","webmcp"],"latest_commit_sha":null,"homepage":"https://las3d-studio.pages.dev","language":"TypeScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/tech-sumit.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-06-18T19:33:27.000Z","updated_at":"2026-07-04T11:13:27.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/tech-sumit/LiveAvatarStream3D","commit_stats":null,"previous_names":["tech-sumit/liveavatarstream3d"],"tags_count":1,"template":false,"template_full_name":null,"purl":"pkg:github/tech-sumit/LiveAvatarStream3D","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tech-sumit%2FLiveAvatarStream3D","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tech-sumit%2FLiveAvatarStream3D/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tech-sumit%2FLiveAvatarStream3D/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tech-sumit%2FLiveAvatarStream3D/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/tech-sumit","download_url":"https://codeload.github.com/tech-sumit/LiveAvatarStream3D/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tech-sumit%2FLiveAvatarStream3D/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35394859,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-12T02:00:06.386Z","response_time":87,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["3d","ai-agents","avatar","cloudflare-workers","elevenlabs","lip-sync","mcp","talking-head","text-to-speech","threejs","typescript","video-generation","virtual-presenter","webcodecs","webmcp"],"created_at":"2026-07-12T15:01:18.452Z","updated_at":"2026-07-12T15:01:19.254Z","avatar_url":"https://github.com/tech-sumit.png","language":"TypeScript","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n# 🎬 LiveAvatarStream3D\n\n**Write a script. A 3D presenter delivers it — lip-synced, camera-directed, and exported to MP4. Entirely in your browser.**\n\n[![CI](https://github.com/tech-sumit/LiveAvatarStream3D/actions/workflows/ci.yml/badge.svg)](https://github.com/tech-sumit/LiveAvatarStream3D/actions/workflows/ci.yml)\n[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)\n[![Node ≥ 20](https://img.shields.io/badge/Node-%E2%89%A5%2020-339933?logo=node.js\u0026logoColor=white)](package.json)\n[![TypeScript](https://img.shields.io/badge/TypeScript-strict-3178C6?logo=typescript\u0026logoColor=white)](tsconfig.base.json)\n[![Three.js](https://img.shields.io/badge/Three.js-WebGL-000000?logo=three.js\u0026logoColor=white)](apps/avatar-live)\n[![PRs Welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](CONTRIBUTING.md)\n\n### [▶ **Try it live** — las3d-studio.pages.dev](https://las3d-studio.pages.dev/)\n\n*No install, no keys: the hosted studio runs the full live preview with free browser voices (Chrome/Edge recommended). Paste your own ElevenLabs key in the Voice panel (it stays in your browser) to unlock real voices and MP4 export right in the demo.*\n\n![LiveAvatarStream3D product demo](docs/media/product-demo.gif)\n\n*Every pixel above was rendered by a browser tab — no render farm, no GPU server, no upload.*\n\n\u003c/div\u003e\n\n---\n\n## Why this exists\n\nTools like HeyGen and Synthesia turn scripts into presenter videos — on their servers, with their avatars, at their price. **LiveAvatarStream3D does it as an open-source studio that runs client-side:**\n\n- 🖥️ **No render server.** The 3D scene renders live with Three.js and exports frame-exact MP4 (1080p or 4K) with WebCodecs, right in the tab.\n- ✍️ **Direction is data, not editing.** Emotion, gestures, and camera work are inline script tags and JSON — versionable, diffable, generatable by an LLM.\n- 🤖 **Agent-native.** The studio exposes its controls as an in-page [WebMCP](https://github.com/webmachinelearning/webmcp) server, so AI agents (Claude Code, or anything MCP-speaking) can author, direct, and export videos autonomously.\n\n## See it\n\nA moment from a newscast produced end-to-end by an AI agent driving this studio — script, camera direction, wall slides, ticker, music bed, and export:\n\n![AI-produced newscast: presenter speaking beside slide wall](docs/media/newscast-finale.gif)\n\nFull 1080p MP4s (with audio) are on the [**Releases page**](https://github.com/tech-sumit/LiveAvatarStream3D/releases) →\n\n\u003e **Note:** all sample newscasts in this repo are **fictional demo content**, written to exercise the pipeline. They are not real news, and this project is not affiliated with or endorsed by any company mentioned in them.\n\n| | |\n|---|---|\n| ![Product pitch two-shot: presenter beside video wall](docs/media/still-pitch.png) | ![Cinematic low-angle hero framing of the presenter](docs/media/still-herolow.png) |\n| The anchor pitches beside the video wall (`two-shot`) | Cinematic low-angle framing (`hero-low`), one tag away |\n\n## The studio\n\n![The studio interface: script editor, voice manager, 3D viewport, export panel](docs/media/studio-ui.png)\n\nOne screen: script editor with emotion/gesture tag palette, voice picker (cloned voices included), live 3D viewport with the exact export framing overlaid, studio lighting rig, and one-click MP4 export.\n\n## Quickstart\n\n**Zero-install:** open the [hosted demo](https://las3d-studio.pages.dev/) — script, direct, and preview entirely in your browser. Local setup unlocks MP4 export with your own ElevenLabs voices:\n\n```bash\ngit clone https://github.com/tech-sumit/LiveAvatarStream3D.git\ncd LiveAvatarStream3D\nnpm install\nbash apps/avatar-live/scripts/fetch-avatars.sh      # avatar models (not redistributed in-repo)\nbash apps/avatar-live/scripts/fetch-animations.sh   # gesture/locomotion clips\nnpm run dev:avatar     # → http://localhost:5175\n```\n\n**Out of the box (no keys):** type a script and get the live lip-synced 3D preview with free browser Web Speech voices.\n**For MP4 export:** add an ElevenLabs key to `apps/avatar-live/.env` (copy [`.env.example`](apps/avatar-live/.env.example)) — the exporter renders from a synthesized narration buffer, and browser Web Speech audio can't be captured.\n\n**Browser support:** Chrome/Edge — full studio + MP4 export. Firefox/Safari — live preview (WebCodecs MP4 support varies; the app falls back to a WebM quick preview). Agent control via WebMCP — Chrome 146+.\n\n\u003e The optional Cloudflare control plane (voice cloning, avatar builds) lives in [`services/`](services/) — the studio runs fully without it. Cloud project storage is a studio-side R2 proxy configured by the `R2_*` keys in the same `.env`; without them, projects persist to localStorage.\n\n## Direct it like a script\n\nPerformance direction lives *inside* the text — emotion and gesture tags, exactly as the parser reads them:\n\n```\n[serious] Good evening. A follow-up tonight on the story we brought you earlier.\n[confident][point] The numbers on the wall tell the story.\n[happy][open_palms] And that changes everything.\n```\n\nFor full productions there's a **newscast document** — sections, beats, camera presets, wall slides, tickers, and a music bed — compiled onto one timing clock with the narration audio. Abridged from a real, valid document ([full schema](packages/protocol/src/newsreport.ts)):\n\n```jsonc\n{\n  \"version\": 2,\n  \"meta\": { \"title\": \"Fable 5 Returns\", \"anchors\": [{ \"id\": \"ava\", \"name\": \"Ava Lin\", \"avatarUrl\": \"avaturn-model\", \"voiceId\": \"...\" }] },\n  \"rundown\": [{\n    \"headline\": \"What Changed\",\n    \"bullets\": [\"Back in Claude Code\", \"A new model tier\"],\n    \"beats\": [{\n      \"text\": \"Here is what changed. Fable five returns as the generally available tier.\",\n      \"gesture\": \"point\",\n      \"camera\": { \"preset\": \"hero-low\" },\n      \"pause_ms_after\": 300\n    }]\n  }]\n}\n```\n\nTen camera framings ship as data (a [catalog table](packages/performer-core/src/cameraShots.ts), not code): `close` · `medium` · `wide` · `two-shot` · `ots-screen` · `profile` · `hero-low` · `dutch` · `establish` · `push-in`.\n\n## Agents can run the whole studio\n\nThe studio registers its controls on the page's model context (the emerging [WebMCP](https://github.com/webmachinelearning/webmcp) standard) — including on the [hosted demo](https://las3d-studio.pages.dev/). Open it in a WebMCP-capable Chrome (146+) and any in-browser LLM or MCP bridge — e.g. [`@tech-sumit/mcp-webmcp`](https://www.npmjs.com/package/@tech-sumit/mcp-webmcp) — can set the newscast, restyle the wall, and export the MP4, hands-free. The demo videos in this README were produced exactly that way, by Claude Code.\n\n**→ [How to connect an agent + the 23-tool catalog](docs/webmcp.md)**\n\n## How it works\n\n```mermaid\nflowchart LR\n  subgraph Browser[\"Browser (the whole render pipeline)\"]\n    Script[\"Script + tags\u003cbr/\u003enewscast JSON\"] --\u003e Compile[\"compile to Score\u003cbr/\u003e(one timing clock)\"]\n    Compile --\u003e Perform[\"performer\u003cbr/\u003eThree.js scene · blendshape lip-sync\u003cbr/\u003egestures · camera presets\"]\n    Perform --\u003e Live[\"live preview (rAF)\"]\n    Perform --\u003e Export[\"offline exporter\u003cbr/\u003eWebCodecs, frame-exact\u003cbr/\u003e1080p / 4K MP4\"]\n    TTS[\"TTS audio\u003cbr/\u003eElevenLabs / Web Speech\"] --\u003e Compile\n  end\n  subgraph CF[\"Cloudflare control plane (optional)\"]\n    API[\"Hono Worker\u003cbr/\u003eD1 · R2 · Queues\"] --\u003e GPU[\"GPU jobs\u003cbr/\u003evoice clone · avatar build\"]\n  end\n  Browser \u003c-. \"voices · avatars\" .-\u003e CF\n```\n\nThe key design decision: **the live preview and the export run the same performance code on the same clock.** What you see is exactly what renders — lip-sync mouth tracks, gesture timing, camera moves, music and SFX are all deterministic per frame.\n\n| Workspace | What it is |\n|---|---|\n| [`apps/avatar-live`](apps/avatar-live) | The studio. Vanilla TS + Three.js; glTF ARKit/viseme blendshape avatars, Mixamo-retargeted gestures/locomotion, WebCodecs export, WebMCP server. |\n| [`packages/performer-core`](packages/performer-core) | Framework-agnostic runtime: the camera shot catalog, pure-math shot composition, motion/gesture solvers. |\n| [`packages/protocol`](packages/protocol) | Shared contracts: the script DSL, the Score model + compiler, the newscast document + compiler, studio bridge tools (zod → JSON Schema). |\n| [`services/control-api`](services/control-api) | Cloudflare Worker (Hono): voices, avatars, uploads, job orchestration; D1 + R2 + Queues + Durable Objects. |\n| [`services/gpu`](services/gpu) | Python GPU jobs the control plane dispatches: voice cloning and avatar builds (plus a standalone image-gen service, not yet wired to the orchestrator). |\n| [`services/newsroom-mcp`](services/newsroom-mcp) | stdio MCP for asset generation: card graphics, montages, music, post-production. |\n\n## Status \u0026 limitations\n\nAn honest map of where this is today:\n\n- **Chromium-first.** MP4 export needs WebCodecs H.264 (Chrome/Edge); other browsers get the live preview and a WebM fallback.\n- **MP4 narration requires an ElevenLabs key.** Web Speech covers the free live preview only — its audio cannot be captured into an export.\n- **Single-user POC.** No auth on the studio; the control plane user is hardcoded. The Worker supports token auth via config.\n- **Export speed:** a 56 s 1080p newscast takes ~8–10 minutes on an M-series MacBook (frame-exact offline render, works in a backgrounded tab).\n- **Tested where it counts:** 380+ unit tests across the workspaces; typecheck + tests in CI on every PR; the demo newscasts here exercised the full pipeline end-to-end.\n\n## Roadmap\n\nThe active direction is making **all** direction data-driven — one interpreter, zero hardcoded performance code:\n\n- **Score/Stage DSL** — a parameterized, spatial, compositional performance model ([design spec](docs/specs/2026-06-25-performance-score-dsl-design.md))\n- **WebMCP-first control** — the studio as a first-class agent tool surface ([design spec](docs/specs/2026-06-25-webmcp-studio-control-design.md))\n- More avatars, more stages, more chrome packs\n\nSpecs live in [`docs/specs/`](docs/specs/) — the repo is built spec-first. The founding brief is preserved in [`docs/original-product-brief.md`](docs/original-product-brief.md).\n\n## Third-party assets\n\nAvatar models and animation clips are **not redistributed in this repo** — the fetch scripts pull them from their sources under their own terms (Avaturn avatars; Mixamo/Ready-Player-Me animation clips). Bundled media credits are listed in [THIRD-PARTY-NOTICES.md](THIRD-PARTY-NOTICES.md). The MIT license covers the **code**; fetched assets remain under their providers' licenses.\n\n## Contributing\n\nIssues and PRs welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). Good first areas: camera presets (it's a data table), gesture clips, stage/chrome styling, docs.\n\n## License\n\n[MIT](LICENSE) © 2026 [Sumit Agrawal](https://sumitagrawal.dev) — code only; third-party assets keep their own licenses.\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\nBuilt by **[Sumit Agrawal](https://sumitagrawal.dev)** ([@tech-sumit](https://github.com/tech-sumit)) — I build agent-native tools, realtime avatar systems, and the occasional robot brain.\n\nIf this saved you a render farm, you know where the star button is.\n\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftech-sumit%2Fliveavatarstream3d","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ftech-sumit%2Fliveavatarstream3d","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftech-sumit%2Fliveavatarstream3d/lists"}