← library
skill🛬 Landing pagesv1 · updated 2026-10-09

scroll-world

Builds a scroll-scrubbed "fly through the world" landing page where scrolling drives a pre-rendered AI camera flight through connected scenes with no visible cuts. Use when asked for a scroll cinematic, diorama landing page, 3D-world hero, or to turn a business into a scrollable world.

Run it as a prompt

Paste this into any AI agent, or fetch it: curl -s https://uplift.page/api/v1/prompts/scroll-world/raw

prompt.md
# scroll-world

Condensed from oso95/scroll-world (MIT licence, notice in `references/LICENSE-scroll-world`). The page scrubs pre-rendered video by scroll position: the camera dives from outside a scene into its interior, then flows into the next scene with no cuts. Step numbers inside `references/` refer to the original 760-line skill; this file keeps the decisions and the traps, and the references keep the runnable material.

**Chain:** N scene stills, then N "dive-in" clips, then N-1 "connector" clips joining consecutive scenes, then a framework-agnostic scrub engine that plays them as one flight.

**The rule that makes or breaks it:** every seam must be frame-identical. Chain each clip from the previous clip's actual rendered frames, never from the original still. Getting this wrong produces a visible pop.

## Before you start

- This spends real money on third-party video generation (roughly $11 to $27 for a six-scene chain, doubled with a mobile version). Never render without a stated estimate and a go from the user.
- The reference pipeline uses the Higgsfield CLI (stills, fallback video), the Monid CLI (default video billing), `ffmpeg`/`ffprobe`, and optionally PIL and the Codex CLI. Check each is installed and authenticated. Authentication is interactive, so ask the user to do it.
- Generations take 3-8 minutes each. Run them detached and poll.
- Run array-driven steps in a `#!/bin/bash` script, not the interactive zsh (arrays are 1-indexed there).

Install it as a skill

Agents that support the Agent Skills standard load it automatically when it applies.

download
curl -fsSL https://uplift.page/p/scroll-world/SKILL.md --create-dirs -o .agents/skills/scroll-world/SKILL.md
or from any MCP client
MCP server: https://uplift.page/mcp
Tool: pull_skill  {"slug": "scroll-world"}

The full skill

Condensed from oso95/scroll-world (MIT licence, notice in references/LICENSE-scroll-world). The page scrubs pre-rendered video by scroll position: the camera dives from outside a scene into its interior, then flows into the next scene with no cuts. Step numbers inside references/ refer to the original 760-line skill; this file keeps the decisions and the traps, and the references keep the runnable material.

Chain: N scene stills, then N "dive-in" clips, then N-1 "connector" clips joining consecutive scenes, then a framework-agnostic scrub engine that plays them as one flight.

The rule that makes or breaks it: every seam must be frame-identical. Chain each clip from the previous clip's actual rendered frames, never from the original still. Getting this wrong produces a visible pop.

Before you start

  • This spends real money on third-party video generation (roughly $11 to $27 for a six-scene chain, doubled with a mobile version). Never render without a stated estimate and a go from the user.
  • The reference pipeline uses the Higgsfield CLI (stills, fallback video), the Monid CLI (default video billing), ffmpeg/ffprobe, and optionally PIL and the Codex CLI. Check each is installed and authenticated. Authentication is interactive, so ask the user to do it.
  • Generations take 3-8 minutes each. Run them detached and poll.
  • Run array-driven steps in a #!/bin/bash script, not the interactive zsh (arrays are 1-indexed there).

1. Interview

Ask the subject as an open question, not a made-up multiple choice. Then settle:

  1. Brand kit: 4-6 named hex values, a display name, a tone. Import from a URL, take it from the user, or propose one.
  2. Art direction: default soft clay diorama, isometric, tilt-shift. The chosen wording becomes one style preamble reused verbatim in every prompt.
  3. Camera style (always ask, it sets the architecture):
    • Fly through the world (architecture B): dive in, pull up and out, hop to the next scene. For miniature or diorama looks.
    • One continuous walkthrough (architecture A): a forward-only take. For grounded or photoreal looks.
    • Locked isometric glide (A plus a fixed-angle clause): calmest and cheapest to re-roll.
  4. Sections: 5-7 scenes from the subject's own value chain. Each has a description, headline, one line of body, 0-3 tags. The last usually carries the CTA.
  5. Mobile: ask explicitly. Yes means a second native 9:16 render (about 2x cost), not a crop of the landscape film.
  6. Budget: estimate N stills + (2N-1) videos (x2 with mobile) + ~15% re-roll headroom. Offer a cheap draft tier first (same model at low resolution), then re-render finals. Calibrate on one still and one clip before extrapolating.

2. Stills

One image per section, all sharing the style preamble, on a plain solid background with no text, 3:2. Use one image source for all N stills, since mixed sources read as style drift. Review them as a set (same angle, palette, light) and re-roll outliers. Optionally knock out the background with references/knockout.py to float the scenes. Keep the stills as posters.

3. Video chain

Use one video model for the whole chain. It must accept a start image, and for connectors also an end image; a model that cannot frame-lock seams is declined, not substituted. Mixed models leave a visible character shift.

  • Architecture A (forward take): generate legs sequentially. Each leg starts from the previous leg's real last frame, has no end image, and the prompt says to keep gliding forward. An end image of a wide shot forces a pull-back, the top cause of stutter. There are no connectors. Every leg ends with a slow forward drift and the next begins by continuing it.
  • Architecture B (dive plus connector): a dive clip per scene from its solid-background still. Each connector uses the last frame of dive i as its start image and the first frame of dive i+1 as its end image, both extracted from the rendered videos. The reversal at each seam only reads as intentional in a miniature world; warn the user before using B on realistic content.
  • Eyeball each leg's last frame before chaining the next. If it is not a clean forward drift, re-roll first.
  • Pick mid-leg camera moves from the concept (orbit for products, steadicam for interiors, lateral track for industrial). A plain forward glide is the zero-risk default.
  • Pass the aspect ratio explicitly (16:9, or 9:16 for mobile). Do not request audio.
  • Frames travel to the Monid backend by public URL, never inline base64.
  • Read the billed cost off every clip. Re-check the backend's input schema before each build, since the catalog changes.

Exact commands, flags and prompt templates are in references/pipeline.md and references/prompts.md.

4. Encode

Encode at native resolution with no downscale: -crf 20 -g 8, audio stripped, +faststart, a light unsharp. Do not use all-intra (files balloon). Mobile encodes are 720 wide, -g 4, crf 23. The engine plays clips from in-memory blobs, which is what makes them seekable on hosts without byte-range support.

5. Assemble

Copy references/scrub-engine.js (and optionally references/index-template.html) into the project and call mountScrollWorld(container, config) with brand, sections[] (each with clip, still, copy, optional clipMobile, stillMobile, scroll, linger) and connectors[]. Theme it with --sw-* CSS variables. The engine already handles scroll-to-time smoothing, lazy prefetch, seam crossfade, reduced motion and phone hardening. Do not hide the still on loadedmetadata or drop playsinline/muted if you port it, or iOS shows blank scenes.

6. QA the seams (do not skip)

  • Screenshot just before and after each seam. Judge by composition, not raw PSNR; a correct seam can read 18-25 dB from detail shimmer alone. A real mismatch shows different props or framing.
  • Confirm no console errors, video.seekable.end(0) > 0, and currentTime tracking scroll.
  • Check reduced motion falls back to the stills.
  • With mobile, test on a throttled phone viewport: fast flick tracks, first scene shows immediately, videoWidth < videoHeight, no page jump when the URL bar collapses. Test iOS Safari specifically.

Common failures

Symptom Cause and fix
Pop at a seam Connector used the still, not the neighbours' real frames. Redo with extracted frames.
Camera seems to jump backward Velocity reverses across a seam (architecture B). Use A for grounded scenes.
Video frozen at frame 0 Host lacks byte ranges. Use blob loading.
Content-filter rejection on an innocent clip Re-roll, strip trigger words and add "empty, unoccupied", then try the alternate model, then leave the connector slot null for a direct crossfade.
Clip is the wrong aspect The adaptive default follows the input image. Set the ratio explicitly.
Phone freezes on a flick Ship the mobile encodes.

References

  • references/prompts.md: intake checklist and every prompt template.
  • references/pipeline.md: copy-paste batch scripts, including the Monid backend wiring.
  • references/scrub-engine.js, references/index-template.html: the engine and a standalone page.
  • references/knockout.py: background knockout.

Served from the uplift.page library and refreshed within 5 minutes of every update.