Three independent production lines: Product Intro / Short Drama Draft / Digital Human Narration — an end-to-end AI video factory from script to final cut
The AI Video Generator is a full-stack video factory covering three independent production lines: Product Intro Videos (Remotion deterministic template rendering), Low-cost Drama Drafts (script → storyboard → per-shot Wan generation → screening), and Digital Human Narration (TTS → face swap → lip sync → denoise subtitle → one-click export). Built on React + FastAPI with LLM copy generation, Edge-TTS speech synthesis, Remotion cinematic rendering, WanGP/face_swap/lip_sync digital human pipeline, and an in-browser refine editor. Supports Web / Electron desktop / PWA mobile, plus GPU cluster management and job queue monitoring backend.






Split "making video" into three specialized production lines — each focused on one scenario, non-interfering; sharing GPU cluster and asset library at the bottom, optimizing UX independently on top.
Product Intro uses Remotion deterministic templates (high-quality reusable); Drama uses script→storyboard→Wan per-shot generation (fast low-cost iteration); Digital Human uses TTS+face swap+lip sync (one-stop narration). All three share the same user system, asset library, and GPU cluster.
LLM auto-generates narration copy → Edge-TTS synthesizes speech (free multi-voice) → ffmpeg/Remotion renders video → auto quality check (duration/coverage/consistency). User only inputs product description or uploads assets; everything else is automatic.
React frontend runs simultaneously as Web browser app, Electron desktop app (native menus/file dialogs/auto-update), and PWA mobile app (add to home screen/offline cache/responsive) — one codebase, three platforms.
Server management dashboard monitors GPU cluster status and job queue in real time. Supports face_swap / lip_sync / wan_draft capability tags; tasks route to matching nodes; heartbeat auto-offlines faulty nodes.
Each production line is deeply optimized for one video scenario, forming a closed loop from input to output.
Remotion deterministic template rendering, high-quality commercial-ready. Supports portrait 1080×1920 / landscape 1280×720, duration 15-90s configurable. LLM generates TTS copy → multi-voice synthesis → BGM auto-match (9 moods→5 tracks) → ASS subtitle burn-in → auto quality check.
Script→storyboard→per-shot Wan generation→draft screening rapid iteration pipeline. Input premise (≤200 chars) → auto storyboard → Wan v12 / local WanGP per-shot render → draft layer fast output / backend refine / auto-filter shots. Ideal for early creative validation.
Full digital human pipeline: TTS → face swap → lip_sync → denoise → BGM → subtitles → export 1080P MP4. Supports custom avatar training (upload reference images), green-screen keying for transparent WebP export, multi-voice switching (1080P/720P).
Remotion-based online video editor requiring zero software installation. Supports 15+ format media import, multi-track timeline editing, canvas property fine-tuning (position/size/rotation/opacity/anchor), precise frame-skip preview, MP4 export.
All three lines' outputs managed centrally in "My Works", tabbed by production line. Each record shows thumbnail/title/status tags (Draft ready/Refined/Completed)/timestamp, with preview/refine/download/batch-template-refine actions.
GPU cluster management (server registration/capability tags/heartbeat detection/online status/offline delete), job queue monitoring (job ID/capability type/linked tasks/queued/completed/error logs), user management, dashboard overview.
Using the Product Intro line as example: full pipeline from product description input to 1080P final video output.
Core indicators from architecture design to user experience.
Not just API calls — every step of video production is engineered, observable, and controllable.
Three lines share auth/asset library/GPU cluster/task queue infrastructure, but each has its own UI flow, render engine, and parameter config — adding new lines does not affect existing ones.
TTS → face swap → lip sync → denoise → subtitle → export fully self-integrated, no dependency on external SaaS platforms — data sovereignty, near-zero marginal cost (Edge-TTS is free).
Standard scenes use ffmpeg fast compositing (second-level output); premium scenes use Remotion cinematic template (high-fps/high-design/shot recipe cards) — choose by need, no wasted compute.
Remotion-based online editor lets users complete import/edit/tune/export without leaving the browser — eliminating the "generate→download→open PR→re-edit→re-export" toolchain breakage for a closed-loop experience.
Convert technical capability into quantifiable, reusable business value.
Product content has been published based on internal materials. The following areas are planned for further development:
Explore Xianma AI solutions in other domains