visual-skills
A collection of prompt generation skills for AI image and video creation, combining professional cinematic dramaturgy to produce copy-paste-ready prompts for major generative AI models.

Suggested prompts
About this skill
Visual Skills
A pair of Skills named video and image. video writes AI-video prompts using the working logic of a director, screenwriter, and editor. image writes still-image prompts using an art director's approach. Both select task-appropriate model syntax and return prompts ready to use.

Core Method
Visual Skills handles dramaturgy before model syntax. For video scenes, it first defines:
- Immediate desire and obstacle
- Spatial relationships among characters and objects
- Where the viewer's eye should land
- Shot rhythm and the reason for each cut
- Filmable environmental details, body actions, and sound or visual anchors for each shot
- One core emotion, motif, object, break, and final image for the piece

Before returning an output, it checks scene structure, shot details, shot function, motivated camera movement, readable geometry, and the five shared anchors. A prompt missing required information does not pass the output gate.
The video Skill
video loads dramaturgy rules, cross-model universal rules, one model-specific syntax file, and task modules for areas such as storyboards, keyframes, genres, and camera direction.
It can produce:
- A single video prompt
- A multi-clip prompt sequence with continuity constraints
- A storyboard table
- A prompt audit with problems, missing elements, and a stronger version
- A director treatment
- Veo JSON
Dedicated model rules include:
| Type | Models |
|---|---|
| Video | Seedance 1.0, 1.5 Pro, 2.0, 2.0 Mini, and 2.5 |
| Video | Kling 1.6–2.6 Pro, 3.0, Turbo, and Omni |
| Video | Veo 3 and 3.1 |
| Image | Nano Banana 2 Lite, 2, and Pro |
| Image | GPT Image 2.5 Flare, 2.5 Sunburst, and legacy 2 |
Runway Gen-4, Luma, Pika, and Sora are covered by the universal video-rules layer.
The image Skill
image supports editorial and product photography, posters, UI mockups, infographics, preservation-critical edits, character continuity across a series, and storyboard or animatic keyframes for the video pipeline.
It selects between Nano Banana and GPT Image according to the task. The source README assigns grounding of real places, extreme aspect ratios, and lower-cost batch work to Nano Banana, while dense text, brand assets, and preservation-critical edits go to GPT Image.
The two Skills can be chained: image creates character sheets and keyframes, then video writes animation prompts based on a motion brief.
What You Provide
Provide the intended scene or image, target duration or aspect ratio, target model, and any existing script, references, character sheets, or prompt to audit. When necessary information is missing, the Skill establishes the required scene and shot constraints before writing the prompt.
Usage Examples
- Write a Seedance prompt for a five-second, multi-shot nighttime kitchen scene.
- Storyboard a 30-second film around one emotion and one anchor object.
- Audit an existing prompt, identify what cannot be executed or is missing, and produce a stronger version.
- Translate a script into a sequence of connected short-video prompts.
- Create product-film keyframes and then write Kling prompts to animate them.
Notes
- Prompts prioritize actions, objects, sounds, lighting, and cut points a camera can record instead of replacing them with abstract adjectives.
- Model coverage and dedicated syntax are updated with the repository; actual generation capabilities still depend on the selected model and its current version.
- This Skill is licensed under CC BY 4.0. Copies and derivatives must retain attribution to Serge Shima.