Condensed from oso95/scroll-world (MIT licence, notice in references/LICENSE-scroll-world). The page scrubs pre-rendered video by scroll position: the camera dives from outside a scene into its interior, then flows into the next scene with no cuts. Step numbers inside references/ refer to the original 760-line skill; this file keeps the decisions and the traps, and the references keep the runnable material.
Chain: N scene stills, then N "dive-in" clips, then N-1 "connector" clips joining consecutive scenes, then a framework-agnostic scrub engine that plays them as one flight.
The rule that makes or breaks it: every seam must be frame-identical. Chain each clip from the previous clip's actual rendered frames, never from the original still. Getting this wrong produces a visible pop.
Before you start
- This spends real money on third-party video generation (roughly $11 to $27 for a six-scene chain, doubled with a mobile version). Never render without a stated estimate and a go from the user.
- The reference pipeline uses the Higgsfield CLI (stills, fallback video), the Monid CLI (default video billing),
ffmpeg/ffprobe, and optionally PIL and the Codex CLI. Check each is installed and authenticated. Authentication is interactive, so ask the user to do it. - Generations take 3-8 minutes each. Run them detached and poll.
- Run array-driven steps in a
#!/bin/bashscript, not the interactive zsh (arrays are 1-indexed there).
1. Interview
Ask the subject as an open question, not a made-up multiple choice. Then settle:
- Brand kit: 4-6 named hex values, a display name, a tone. Import from a URL, take it from the user, or propose one.
- Art direction: default soft clay diorama, isometric, tilt-shift. The chosen wording becomes one style preamble reused verbatim in every prompt.
- Camera style (always ask, it sets the architecture):
- Fly through the world (architecture B): dive in, pull up and out, hop to the next scene. For miniature or diorama looks.
- One continuous walkthrough (architecture A): a forward-only take. For grounded or photoreal looks.
- Locked isometric glide (A plus a fixed-angle clause): calmest and cheapest to re-roll.
- Sections: 5-7 scenes from the subject's own value chain. Each has a description, headline, one line of body, 0-3 tags. The last usually carries the CTA.
- Mobile: ask explicitly. Yes means a second native 9:16 render (about 2x cost), not a crop of the landscape film.
- Budget: estimate
N stills + (2N-1) videos (x2 with mobile) + ~15% re-roll headroom. Offer a cheap draft tier first (same model at low resolution), then re-render finals. Calibrate on one still and one clip before extrapolating.
2. Stills
One image per section, all sharing the style preamble, on a plain solid background with no text, 3:2. Use one image source for all N stills, since mixed sources read as style drift. Review them as a set (same angle, palette, light) and re-roll outliers. Optionally knock out the background with references/knockout.py to float the scenes. Keep the stills as posters.
3. Video chain
Use one video model for the whole chain. It must accept a start image, and for connectors also an end image; a model that cannot frame-lock seams is declined, not substituted. Mixed models leave a visible character shift.
- Architecture A (forward take): generate legs sequentially. Each leg starts from the previous leg's real last frame, has no end image, and the prompt says to keep gliding forward. An end image of a wide shot forces a pull-back, the top cause of stutter. There are no connectors. Every leg ends with a slow forward drift and the next begins by continuing it.
- Architecture B (dive plus connector): a dive clip per scene from its solid-background still. Each connector uses the last frame of dive i as its start image and the first frame of dive i+1 as its end image, both extracted from the rendered videos. The reversal at each seam only reads as intentional in a miniature world; warn the user before using B on realistic content.
- Eyeball each leg's last frame before chaining the next. If it is not a clean forward drift, re-roll first.
- Pick mid-leg camera moves from the concept (orbit for products, steadicam for interiors, lateral track for industrial). A plain forward glide is the zero-risk default.
- Pass the aspect ratio explicitly (
16:9, or9:16for mobile). Do not request audio. - Frames travel to the Monid backend by public URL, never inline base64.
- Read the billed cost off every clip. Re-check the backend's input schema before each build, since the catalog changes.
Exact commands, flags and prompt templates are in references/pipeline.md and references/prompts.md.
4. Encode
Encode at native resolution with no downscale: -crf 20 -g 8, audio stripped, +faststart, a light unsharp. Do not use all-intra (files balloon). Mobile encodes are 720 wide, -g 4, crf 23. The engine plays clips from in-memory blobs, which is what makes them seekable on hosts without byte-range support.
5. Assemble
Copy references/scrub-engine.js (and optionally references/index-template.html) into the project and call mountScrollWorld(container, config) with brand, sections[] (each with clip, still, copy, optional clipMobile, stillMobile, scroll, linger) and connectors[]. Theme it with --sw-* CSS variables. The engine already handles scroll-to-time smoothing, lazy prefetch, seam crossfade, reduced motion and phone hardening. Do not hide the still on loadedmetadata or drop playsinline/muted if you port it, or iOS shows blank scenes.
6. QA the seams (do not skip)
- Screenshot just before and after each seam. Judge by composition, not raw PSNR; a correct seam can read 18-25 dB from detail shimmer alone. A real mismatch shows different props or framing.
- Confirm no console errors,
video.seekable.end(0) > 0, andcurrentTimetracking scroll. - Check reduced motion falls back to the stills.
- With mobile, test on a throttled phone viewport: fast flick tracks, first scene shows immediately,
videoWidth < videoHeight, no page jump when the URL bar collapses. Test iOS Safari specifically.
Common failures
| Symptom | Cause and fix |
|---|---|
| Pop at a seam | Connector used the still, not the neighbours' real frames. Redo with extracted frames. |
| Camera seems to jump backward | Velocity reverses across a seam (architecture B). Use A for grounded scenes. |
| Video frozen at frame 0 | Host lacks byte ranges. Use blob loading. |
| Content-filter rejection on an innocent clip | Re-roll, strip trigger words and add "empty, unoccupied", then try the alternate model, then leave the connector slot null for a direct crossfade. |
| Clip is the wrong aspect | The adaptive default follows the input image. Set the ratio explicitly. |
| Phone freezes on a flick | Ship the mobile encodes. |
References
references/prompts.md: intake checklist and every prompt template.references/pipeline.md: copy-paste batch scripts, including the Monid backend wiring.references/scrub-engine.js,references/index-template.html: the engine and a standalone page.references/knockout.py: background knockout.