Screenplay PDF
Industry-format LaTeX screenplay compiled from Fountain-structured scenes — ready for readers and producers.
A hierarchical pipeline of local AI agents turns a PDF novel into industry-format screenplay artifacts and a 4K picture-video — entirely on your machine, with no cloud APIs.
Architecture
Specialized local agents handle global structure, parallel scene writing, and multimedia synthesis — with checkpointing at every step.
PDF → structured text chunks
Story bible & 3-act beat sheet
~70 scene skeletons with dramatic goals
Parallel Fountain-structured scenes
Cross-scene consistency patches
Dialogue & action refinement
Fountain, LaTeX & PDF output
Cinematic still per scene + portraits
Per-character voice synthesis
Ken Burns, captions & chapters
Deliverables
Every run yields a complete adaptation package — from formatted screenplay to a watchable narrated film, without sending your manuscript to the cloud.
Industry-format LaTeX screenplay compiled from Fountain-structured scenes — ready for readers and producers.
Characters with voice & arc, locations, themes, and a 24–40 beat three-act treatment distilled from the source book.
One cinematic 1920×1080 still per scene, plus character portrait thumbnails for rich-caption mode.
Per-line WAV tracks with distinct macOS voices assigned to each speaking character in reading order.
3840×2160 MP4 with Ken Burns motion, per-character coloured captions, slug cards, act breaks, and chapter metadata.
SRT file emitted alongside every render — usable independently of on-video overlay presets.
Adaptation Draft
Int. Study — Night
Rain hammers the window. A manuscript lies open, pages curling at the edges. The lamp throws a warm cone of light across the desk.
Protagonist
(quietly, to the empty room)
Every story wants to become something else.
They close the book. On the monitor, a progress bar inches forward.
· · · Generated scene — illustrative preview · · ·
Operational Benchmarks
Observed on a 96 GB Mac Studio. End-to-end wall-clock for a full two-hour pipeline: several hours — dominated by LLM generation and 4K ffmpeg encodes.
Target runtime
120 min
≈120 script pages, ~70 scenes
Peak memory
~48 GB
96 GB Mac Studio, throttled cap
Text model
14B params
Director, planner, continuity
Writer model
8B params
3 parallel scene agents
Beats
24–40
Three-act structure
Disk (2 h run)
≥30 GB
Intermediates + final outputs
| Ollama instances | Peak RAM | Notes |
|---|---|---|
| 1 | ~30 GB | Single server, model swapping |
| 2default | ~48 GB | Default — text + image servers |
| 4 | ~80 GB | Higher writer throughput |
Key Findings
14B models handle global structure; 8B writers run in parallel with half the KV-cache footprint per slot.
Text models are evicted before image generation; image models before ffmpeg — reclaiming ~10 GB between stages.
Every stage caches under work/. Crashes and Ctrl-C are safe; re-runs skip completed steps.
Lenient parsing with auto-close handles truncated LLM output — critical for long scene-plan responses.
Parallel writers drift; the continuity pass fixes naming and anachronisms but cannot invent plot facts absent from the bible.
macOS-only TTS today; Ollama image API shapes still evolving; voice availability varies by install.
Research Paper
A hierarchical multi-agent pipeline for local book-to-screenplay and narrated video adaptation — architecture, benchmarks, and future work.
Download PDF