← Back to All Products
Product Plan · AI Video Generation · Content ProductionProduct Plan · Internal Prototype Showcase
🎬

AI Video Generator · All-in-One Video Production Platform

Three independent production lines: Product Intro / Short Drama Draft / Digital Human Narration — an end-to-end AI video factory from script to final cut

The AI Video Generator is a full-stack video factory covering three independent production lines: Product Intro Videos (Remotion deterministic template rendering), Low-cost Drama Drafts (script → storyboard → per-shot Wan generation → screening), and Digital Human Narration (TTS → face swap → lip sync → denoise subtitle → one-click export). Built on React + FastAPI with LLM copy generation, Edge-TTS speech synthesis, Remotion cinematic rendering, WanGP/face_swap/lip_sync digital human pipeline, and an in-browser refine editor. Supports Web / Electron desktop / PWA mobile, plus GPU cluster management and job queue monitoring backend.

Product Demo

Core Modules

🖥️

Choose a Production Line

PRODUCT MODULE
Choose a Production Line
🖥️

Drama Draft Studio

PRODUCT MODULE
Drama Draft Studio
🖥️

Digital Human Studio

PRODUCT MODULE
Digital Human Studio
🖥️

Unified Works Management

PRODUCT MODULE
Unified Works Management
🖥️

In-Browser Refine Editor

PRODUCT MODULE
In-Browser Refine Editor
🖥️

GPU Cluster & Job Management

PRODUCT MODULE
GPU Cluster & Job Management

Design Philosophy

Split "making video" into three specialized production lines — each focused on one scenario, non-interfering; sharing GPU cluster and asset library at the bottom, optimizing UX independently on top.

🏭

Three Parallel Lines, Each Specialized

THREE_LINES

Product Intro uses Remotion deterministic templates (high-quality reusable); Drama uses script→storyboard→Wan per-shot generation (fast low-cost iteration); Digital Human uses TTS+face swap+lip sync (one-stop narration). All three share the same user system, asset library, and GPU cluster.

🔄

Copy → Video Full Pipeline Automation

FULL_PIPELINE

LLM auto-generates narration copy → Edge-TTS synthesizes speech (free multi-voice) → ffmpeg/Remotion renders video → auto quality check (duration/coverage/consistency). User only inputs product description or uploads assets; everything else is automatic.

🖥️

Web + Desktop + Mobile

MULTI_PLATFORM

React frontend runs simultaneously as Web browser app, Electron desktop app (native menus/file dialogs/auto-update), and PWA mobile app (add to home screen/offline cache/responsive) — one codebase, three platforms.

⚙️

GPU Cluster + Task Queue

GPU_CLUSTER

Server management dashboard monitors GPU cluster status and job queue in real time. Supports face_swap / lip_sync / wan_draft capability tags; tasks route to matching nodes; heartbeat auto-offlines faulty nodes.

Core Capabilities · Three Production Lines Explained

Each production line is deeply optimized for one video scenario, forming a closed loop from input to output.

🎬

Product Intro Video Line

Remotion deterministic template rendering, high-quality commercial-ready. Supports portrait 1080×1920 / landscape 1280×720, duration 15-90s configurable. LLM generates TTS copy → multi-voice synthesis → BGM auto-match (9 moods→5 tracks) → ASS subtitle burn-in → auto quality check.

🎭

Short Drama Draft Line

Script→storyboard→per-shot Wan generation→draft screening rapid iteration pipeline. Input premise (≤200 chars) → auto storyboard → Wan v12 / local WanGP per-shot render → draft layer fast output / backend refine / auto-filter shots. Ideal for early creative validation.

👤

Digital Human Narration Line

Full digital human pipeline: TTS → face swap → lip_sync → denoise → BGM → subtitles → export 1080P MP4. Supports custom avatar training (upload reference images), green-screen keying for transparent WebP export, multi-voice switching (1080P/720P).

✏️

In-Browser Refine Editor

Remotion-based online video editor requiring zero software installation. Supports 15+ format media import, multi-track timeline editing, canvas property fine-tuning (position/size/rotation/opacity/anchor), precise frame-skip preview, MP4 export.

📋

Unified Works Hub

All three lines' outputs managed centrally in "My Works", tabbed by production line. Each record shows thumbnail/title/status tags (Draft ready/Refined/Completed)/timestamp, with preview/refine/download/batch-template-refine actions.

🔧

Server Management & Ops

GPU cluster management (server registration/capability tags/heartbeat detection/online status/offline delete), job queue monitoring (job ID/capability type/linked tasks/queued/completed/error logs), user management, dashboard overview.

Product Intro Video Generation Pipeline

Using the Product Intro line as example: full pipeline from product description input to 1080P final video output.

Engineering Flow
1
Input Product Info
Fill in product name, description, target duration (15/30/45/60/90s), orientation (portrait/landscape).
2
LLM Generates Copy
DeepSeek/Kimi/OpenAI-compatible API auto-generates structured narration copy; falls back to placeholder on failure.
3
TTS Speech Synthesis
Edge-TTS free multi-voice synthesis, outputs MP3 + WebVTT subtitle files.
4
BGM + Subtitle Match
Auto-matches BGM mood by video type (9 moods→5 track library); VTT subtitles burned in via ASS.
5
Remotion Render
Cinematic mode uses Remotion high-fps high-design template; standard mode uses ffmpeg fast compositing.
6
Quality Check
Auto-checks video duration, narration coverage rate, product consistency (optional reference image comparison).
7
Preview / Refine / Download
Preview final output in My Works; open browser editor for refinement when needed; download MP4.

Key Metrics

Core indicators from architecture design to user experience.

3
Production Lines
Product Intro / Drama Draft / Digital Human
6+
GPU Capability Tags
face_swap / lip_sync / wan_draft etc.
85
Backend Test Cases
Covers core modules / quality checks / render pipeline
3
Platforms
Web / Electron / PWA

Technical Breakthroughs

Not just API calls — every step of video production is engineered, observable, and controllable.

🔀

Shared Bottom Layer, Independent Top Layers

Three lines share auth/asset library/GPU cluster/task queue infrastructure, but each has its own UI flow, render engine, and parameter config — adding new lines does not affect existing ones.

🤖

Self-Built Full Digital Human Pipeline

TTS → face swap → lip sync → denoise → subtitle → export fully self-integrated, no dependency on external SaaS platforms — data sovereignty, near-zero marginal cost (Edge-TTS is free).

🎞️

Dual Render Engine Adaptation

Standard scenes use ffmpeg fast compositing (second-level output); premium scenes use Remotion cinematic template (high-fps/high-design/shot recipe cards) — choose by need, no wasted compute.

🌐

In-Browser Post-Production

Remotion-based online editor lets users complete import/edit/tune/export without leaving the browser — eliminating the "generate→download→open PR→re-edit→re-export" toolchain breakage for a closed-loop experience.

Business Value

Convert technical capability into quantifiable, reusable business value.

Improvements Delivered
Consolidates three common video needs (product video/drama/narration) into one platform — no more switching between tools.
Drama line compresses traditional 1-week animation storyboard iteration to hours via "draft layer fast output" — creative validation cost down 90%+.
Fully self-built digital human pipeline (TTS+swap+lip sync) costs nearly zero marginal expense vs. thousands/month for commercial digital human SaaS.
GPU cluster management enables multi-machine multi-GPU collaboration — face_swap/lip_sync/wan_draft capabilities route to different GPU nodes.
In-browser refine editor eliminates toolchain fracture: no more generate→download→open Premiere/Cut→edit→re-export loop.
Web + Electron + PWA triple coverage: team uses Web in office, mobile on the go, pros install desktop — one codebase serves all scenarios.
Applicable Scenarios
E-commerce product intro videos: enter product link/description → auto-generate 30-60s subtitled promo video for detail pages / social media / TikTok.
Short drama script rapid validation: screenwriter writes premise → drama line auto-storyboards+Wan renders draft → team screens and decides on full production investment.
Corporate training/knowledge narration: trainer uploads photo → trains digital human → inputs lecture script → one-click generates narration video for internal training or public courses.
Livestream clip post-processing: import recording into refine editor → trim highlights → add subtitles/BGM → export for distribution across platforms.
📌

Content Under Active Update

Product content has been published based on internal materials. The following areas are planned for further development:

○
More Remotion cinematic templates (currently product intros; expand to brand stories/event recaps/data reports etc.)
○
Digital human avatar marketplace (UGC upload/share/rent custom avatars) and richer expression/action repertoire
○
Deeper drama line integration with mainstream AI video models (Sora/Kling/Kling etc. as Wan alternative backends)
○
Batch production & A/B testing (same copy × different styles/hosts × different BGMs = N-version matrix)
○
Collaboration mode (multi-user simultaneous edit/review/comment/version comparison) and refined permission system
○
CDN distribution & extended i18n (current ZH/EN UI; expand to JA/KO/ES/AR etc.)
Contact Us

Start AI Partnership

Whether in government, finance, manufacturing, consumer, or content, we can customize vertical AI agent solutions for you.

📍

Address

Xiamen, Fujian · Wuhan OPC (planned)

🌐
贤

Xianma AI

Xiamen Xianma Intelligent Technology Co., Ltd.

© 2024-2026 Xiamen Xianma Intelligent Technology Co., Ltd. · AI Agent Solutions · www.xianma.top

Products: 19active projects