SKILL · Codex 插件存档 · pixverse

references/source-preserving.md

This workflow accepts exactly one measured 4–30s source video. Preserve that full duration, display aspect, motion, performance, camera, cuts, lighting, pacing and default audio. An out-of-range source needs a genuinely supported broader route or an explicit change of scope, not a secret trim, loop, freeze or clamp. Inspect current PixVerse modify/edit capabilities before spending: generative reference video is not proof of exact video editing.

One Mapping Per Output

Resolve the requested ordered output list or count, source targets, timing and replacement references. Zip ordered lists; never create a Cartesian product. Each output starts from the original source, not the previous variant. Two missing adult replacement people in each of N outputs means 2N distinct reference people, not N or a reused pair by accident. Only generate missing references actually needed by the edit; removal needs none.

Reuse usable supplied or user-authorized reference images before creating new assets. For a missing adult replacement, use the shared Sunburst 2K/high policy. Make one clear reference per distinct person: eye-level, straight-on, full body, one adult, direct gaze, concrete hair and complete opaque wardrobe/footwear, neutral matte-white studio, soft even light, natural skin and well-resolved hands. Exclude unrelated props, logos and collage panels. Distinguish replacements through chosen hair, clothing and other visible styling; do not derive sensitive demographics from source appearance. Check the resulting reference before spending on its dependent video. Retain the user's authority to choose or reject people; when reference preparation is delegated, proceed with a suitable candidate under the existing generation-confirmation policy.

Inspect the source once, retaining duration/streams/aspect, timed scenes, target appearances, cuts and visible text. Reuse analysis; retry a failed/unusable analysis only once before resolving missing evidence. Describe visible hairstyle/clothing/pose and use supplied casting choices; do not infer ethnicity or other sensitive identity from a person's image. Source-platform demographic-contrast requirements are not needed to make identities distinct.

Map references by explicit user mapping, declared role, unique visible/functional match, then stable supplied order only when unambiguous. A character sheet is one person; several views may be complementary inputs. Do not silently discard a reference. A person image normally owns their complete visible look: face, hair, build, clothes, footwear and worn accessories. A separate garment/full-outfit reference or explicit wardrobe instruction can override clothing. Source held/environmental props remain unless separately targeted.

A garment board can contain a bag, eyewear or jewelry that conflicts with protected ad content. State the selected garments and excluded accessories explicitly before prompt writing. A separate protected-product view is complementary identity evidence; it does not request a second product, new cutaway or inserted collage. A location reference controls room geometry/materials, while the source controls viewpoints and performance; map supporting furniture where the performer sits or leans rather than freezing the reference photograph behind a moving camera.

Identity replacement normally follows every appearance, including occlusion, reflected views, shadows, blur and transitions; preserve all unmapped people. If the user explicitly requests a partial/time-scoped identity effect, respect it and confirm the selected route can express that scope; never widen the edit silently. Attribute/clothing changes keep the same person and do not trigger whole-person exclusion.

Operation Grammar

Use actual source timing: an exact user range wins, otherwise map the event to an observed boundary or the target's visible range. Do not fragment around momentary occlusion. Scene-analysis prose can group speech beats or round cut times. Verify any explicit cut schedule against the footage; do not invent cuts from analysis labels. Whole-clip edits can inherit the source cuts without re-describing every gesture or imposing new timings.

  • Replace: named target, actual reference and complete transferred appearance.
  • Modify: only the named property, with all other properties preserved.
  • Remove: reconstruct only the revealed area from its immediate surroundings.
  • Add: placement, size, timing, motion and interaction grounded in the source.
  • Text/graphic edit: exact target, wording and requested style/placement/animation/window;

keep every untargeted property and all other text unchanged.

Every prompt must explicitly preserve all untargeted captions, subtitles, UI, labels, branding and graphics, including wording, style, placement, animation and timing. Text physically attached to a replaced object follows that replacement. Do not automatically remove or regenerate subtitles. A silent generation instruction governs temporary audio, not permission to deliver a silent replacement for an audible source.

Use a compact prompt for a simple non-identity edit: tile the full source duration with unchanged/changed segments, no gaps or overlaps, merging adjacent identical states. Use a detailed prompt for identity, multi-shot interaction, several edit categories or explicit detail: ordered reference declarations → text-preservation block → source-continuity scope → one numbered target/operation at a time → identity/reference locks → resolved render instruction → untouched-content protection. Exclude each mapped original identity where replaced and prevent replacement identities from crossing, merging or duplicating.

Resolve all branches before submission. Check every requested operation/reference appears, no undeclared tag/URL/job ID leaks into prose, and no unrelated camera/scene change is invented. The source prompt ceiling is 3900 characters; retain concise precision and respect any stricter actual model limit. Use PixVerse's actual reference syntax/order, not unverified source-platform tags or payloads.

Before writing an executable edit prompt, read edit-prompt.md for the compact/detailed forms and the reference-precedence check. Keep one complete-look lock per replaced identity; a short generic background-edit example is insufficient for a recast.

Audio, Measurement And Delivery

Where supported, generate visual edits without a new soundtrack, then restore the source's default audio deterministically. If native silence is unsupported, discard only the raw edit's audio. A silent source stays silent. Do not replace speech with TTS or add music. Keep exact source duration through a valid trim/remux only when the full action survives; a generated timing drift, missing cut or changed lips must be reported/repaired, not hidden by muxing the original track. Never freeze/stretch a deficient edit to pretend preservation.

Read media-finalization.md for default-stream selection, timestamp offsets, measured tolerances and a source-audio restoration command. A prompt requesting silence or preserved audio is not evidence that the exported track meets that requirement.

Check every output against source at seams, difficult actions, occlusion/reflection and text moments; probe exact duration, display aspect, chosen resolution and final audio. Deliver only the final source-audio-preserving results in original order; raw silent edits are intermediates. Preserve successful indices; retry only failed ones within the existing approval/test authorization. Report completed/pending/failed counts and specific preservation limitations. Prompt-only work includes the source change sheet and complete per-output prompts.