AI-Powered Script, Voice & Video Engine
SudoVid parses raw production notes and generates fully packaged, ready-to-publish video assets complete with targeted research background frameworks, multi-voice neural audios, and dynamically sequenced video containers.
Requirements
- ▸ Python 3.11+ build runtime profile.
- ▸
ffmpegsystem layer access (automatically managed via environmental orchestration layers on Cloud deployments).
Local Installation
Clone and setup execution instances locally using standard package managers:
# Clone source repository
git clone https://github.com/sudodocs/sudovid.git
cd sudovid
# Install environment requirements
pip install -r requirements.txt
# Start local server
streamlit run app.py
Deploying to Streamlit Cloud
Ensure your deployment repository root contains these specific orchestration configs:
Note: Yourpackages.txtfile must contain a single entry line readingffmpegto force binary installations during deployment environments.
API Credentials Mapping
Enter these verification tokens within the UI tab blocks during data synthesis steps:
| Credential Key | Acquisition Origin | Pipeline Consumption Scope |
|---|---|---|
| Gemini API Key | aistudio.google.com | Research mapping, Script structures, Metadata generation. |
| TMDB API Key | themoviedb.org/settings/api | Film backdrops, official promotional art assets, and cast photos. |
| Pexels API Key | pixels.com/api | Atmospheric stock video assets and general context B-roll clips. |
Content Execution Modes
🎬 Film & Series
Leverages official video trailer extracts via yt-dlp pipelines, structural movie art assets, and cast metadata processing vectors.
🔍 Tech News
Prioritizes contextual verified Wikipedia media layers, B-roll structures, and generates sharp impact data visualizer metric frames.
📚 Educational Tech
Outputs syntax-colored development code frames, structural card break timelines, and explanatory visual parameters.
The 6-Tab Processing Loop
Parameters Configuration
Set video slants, lengths, and custom perspectives via tone matrix adjusters.
Targeted Ground Research
Executes deep context lookups with search engines enabled to parse structural source facts.
Script Synthesis Engine
Generates editable conversational markdown scripts comprising clear hooks, acts, and outcomes.
Neural Audio Generation
Transforms scripts into crisp vocals using high-fidelity edge network reading sound nodes.
Metadata Bundle Assembly
Generates high-CTR descriptions, SEO targeting tag layers, and thumbnail prompts.
Compositing & MP4 Renders
Fetches visual layers, builds timed shot blocks with fluid zoom parameters, overlays text scripts, and renders out structural MP4 file containers.
Format Output Canvas Specifications
- Dimensions: 1080 × 1920 (9:16 aspect canvas)
- Timing Parameters: Under 60 seconds duration limit
- Overlays: Pill-encased centralized header titles
- Dimensions: 1280 × 720 (16:9 aspect canvas)
- Timing Parameters: Matches exact parsed text runtime
- Overlays: Lower-left minimal opacity tags
Known System Boundaries
Execution Safety Parameters
• Keep the web runtime interface tab open continuously while video compilation rendering scripts execute.
• Interface states flush completely during manual browser reload events—always clear tasks cleanly using UI control buttons inside side bars.
• Free deployment execution pipelines possess shared basic computation thresholds, rendering deeper long-form analyses across slower multi-minute pipelines.