Video SuperPowers
Introduction to the video generation UI
This episode introduces the Video Generator interface inside the Magnific AI Suite.
You’ll explore how video generation is organized, how different video models are presented, and what information each model provides before generating anything. The focus is on understanding the layout, available options, and how the interface adapts depending on model capabilities.
By the end of this episode, you’ll feel comfortable navigating the Video Generator, making it easier to focus on prompts and creative decisions in the episodes that follow.
Prompting for video
This episode introduces how to write effective prompts for video generation.
You’ll learn how video prompts differ from image prompts, and why describing movement, camera behavior, and progression over time is essential for stable results. The episode focuses on structuring prompts so the model understands not only what appears in the scene, but how the video should evolve from start to finish.
By the end of this episode, you’ll understand how to write clear, intentional video prompts that produce more controlled and predictable motion.
Using Start and End image for video generation
This episode shows how Start Image and End Image can be used to control a video from the first frame to the last.
You’ll learn how a video can begin from a specific image, progress through a defined action, and resolve into a precise final frame. The episode demonstrates how prompts describe what happens in between, while images anchor the start and the outcome.
By the end of this episode, you’ll understand how to guide video sequences with greater structure, continuity, and predictability using visual anchors.
Using reference video to control movement
This episode shows how video references can be used to control movement in video generation.
You’ll learn how an existing motion sequence can guide the way a character moves, independent of how that character looks. The episode focuses on transferring rhythm, weight, and physical behavior from a reference video to a generated figure.
By the end of this episode, you’ll understand how video references help achieve more natural, coherent motion and reduce unpredictability in animated video results.
Keeping character consistency using image references
This episode focuses on maintaining character consistency throughout a video using image references.
You’ll learn how reference images help preserve a character’s identity as a scene unfolds, even when movement, camera motion, and environment change. The episode shows how visual references stabilize facial features, proportions, and overall appearance over time.
By the end of this episode, you’ll understand how to use image references to keep characters recognizable and consistent across an entire video sequence.
Create your tracks using Music Generator
This episode introduces how to create music using the Music Generator.
You’ll learn how to guide music generation through clear, intentional prompts, and how different levels of structure affect the final result. The episode shows how music prompts can range from simple descriptions to more detailed briefs, depending on whether vocals, lyrics, or specific musical behavior are required.
By the end of this episode, you’ll understand how to approach music generation with clarity, and how to write prompts that produce more predictable and usable tracks.
Understanding Voiceover
This episode explores how written scripts are transformed into voiceovers using the Voice Generator.
You’ll learn how different voice models behave, how tone and delivery can be guided, and how single narration and multi-speaker dialogue are handled. The episode focuses on understanding control rather than complexity.
By the end of this episode, you’ll know how to approach voiceover creation with intention, choosing the right setup depending on whether you need narration, emotion, or conversation.
Get familiar with sound effects
This episode introduces how to create sound effects using the Sound Effects Generator.
You’ll learn how short, direct prompts are used to generate usable sound effects, and how duration, looping, and variation settings influence the final result. The focus is on simplicity and clarity rather than complex descriptions.
By the end of this episode, you’ll understand how to generate practical sound effects that can be easily integrated into videos, animations, or other creative projects.
Mastering Templates
This episode brings together everything you’ve learned throughout the course through the use of templates.
You’ll explore how templates work across image, video, and audio generation, and how they capture complete setups including prompts, references, and settings. The episode shows how templates help you move faster, stay consistent, and avoid starting from scratch every time.
By the end of this episode, you’ll understand how to use and create templates as reusable starting points, turning individual creations into reliable workflows.
From references to a complete workflow
This episode brings together everything learned throughout the course into one complete connected workflow inside Spaces.
Starting from visual references, you’ll build a full promotional piece step by step: generating characters, products, and environments, creating profile sheets for consistency, developing still images, transforming them into video clips, trimming and combining sequences, upscaling the final edit, and adding music to complete the project.
By the end of this episode, you’ll understand how separate Nodes and workflows can work together as one unified creative system from concept to final delivery.
Extracting frames from video
This episode explores how videos can be broken into reusable clips, frames, and audio inside Spaces workflows using Media Extractor.
You’ll learn how to isolate video segments, extract individual frames, sample frames at custom intervals, and separate audio directly from existing clips. The episode also demonstrates how extracted frames can immediately continue into new video workflows, allowing one video to become the starting point for another.
By the end of this episode, you’ll understand how Media Extractor transforms finished videos into reusable workflow assets instead of treating them as final outputs.
Combining and upscaling videos
This episode explores how separate video clips can be combined, refined, and prepared for final delivery inside Spaces workflows.
You’ll learn how Video Combiner merges multiple clips into a continuous sequence, how clip order and transitions are managed inside the workflow, and how Video Upscaler improves final resolution while preserving visual consistency. The episode also demonstrates how audio can be integrated directly into the combined result.
By the end of this episode, you’ll understand how generated clips can continue evolving after creation, becoming part of a larger post-generation workflow inside Spaces.
Connecting images to video generation
This episode explores how image workflows can evolve into video workflows inside Spaces.
You’ll learn how separate visual elements such as characters, environments, and prompts can connect to a Video Generator to create structured motion sequences. The episode also explains how references, Start Images, and text prompts work together to guide video generation while keeping workflows modular and reusable.
By the end of this episode, you’ll understand how video generation can grow naturally from image-based workflows instead of starting from isolated prompts.
Producing UGC-style videos in Spaces
This tutorial shows how to create UGC-style videos in Spaces by building a consistent digital influencer and placing her in different environments with integrated products.
You start by generating a character using visual references and the Assistant, then expand it into a full profile sheet to ensure consistency across views. The workflow continues by creating scene setups, placing the character into environments, and introducing products for review-style content.
Finally, the process moves into video generation, where prompts define the character’s behavior while maintaining strong visual consistency through references. The tutorial demonstrates how to structure a reusable workflow that scales across multiple scenes and products.
By the end, you understand how to combine image generation, references, and video workflows inside Spaces to produce consistent UGC-style content efficiently.
Motion Transfer with Kling Motion Control
This tutorial shows how to use Kling 2.6 Motion Control to transfer real movement from a reference video onto a completely different character.
The workflow separates motion from identity: the video defines the performance, while the image defines who performs it. Using a boxer as the motion source and a mythological bronze statue as the character, the tutorial explains how motion transfer works in practice and why the prompt should focus only on environment, lighting, and camera behavior rather than describing the action again.
It also covers key settings such as resolution limits and Scene Source, showing how aspect ratio can follow either the reference video or the character image. In a second stage, the process expands into the Image Generator, where the character’s environment is redesigned before the same motion is transferred again in Kling 2.6 Motion Control.
By the end, the tutorial makes clear how to separate motion, identity, and environment into distinct steps, making it possible to reuse the same performance across different characters and scenes.
Prompts:
Greek god, blue sky and boxer:
Transform the scene into an ultra-realistic cinematic image of a mythological bronze statue standing centrally, framed from mid-thigh upward with generous headroom above the hair and full visibility of both hands. Remove the Pantheon structure behind the statue entirely, eliminating all architectural clutter in the center background. Keep only two large classical marble columns positioned symmetrically on the far left and far right edges of the frame, slightly out of focus to create depth. The entire central background should now be an expansive bright sky with soft white clouds and subtle atmospheric haze, creating a clean, open, epic backdrop. The sky should be luminous and airy, not dramatic or stormy, with natural daylight clarity and soft gradients of blue. The bronze statue remains the clear focal point, sharply detailed with realistic patina, warm highlights, darker oxidized recesses, and subtle green tones in creases. Lighting is bright, natural Mediterranean daylight with soft directional shadows sculpting the musculature while preserving fine metallic texture. The composition is perfectly centered and balanced, minimal visual noise, strong monumentality, elegant negative space above the head, and cinematic realism. The columns frame the statue without competing with it, and the open sky enhances scale and clarity. High dynamic range, crisp detail, refined contrast, and premium historical film aesthetic.
Change of background greek god in nano banana:
Change the background of my @char1 to an ultra-cinematic professional boxing ring environment at night. The ring is dramatic, high-end, and visually stunning, with powerful overhead stage lighting rigs casting defined beams of light through subtle atmospheric haze. The ropes are perfectly detailed, canvas slightly textured, corners clean and realistic. The lighting is intense but controlled, with warm spotlights from above and soft shadow gradients across the floor. The arena feels large and immersive but not overcrowded, with darkened surroundings and faint audience silhouettes in the distance for depth. Every detail is carefully refined: realistic reflections on metal structures, subtle dust particles in the light beams, balanced contrast, premium color grading. The background must feel epic and cinematic but never overpower my character. My character remains perfectly integrated into the environment with consistent lighting direction, realistic shadow contact with the ring floor, correct perspective scale, and natural depth of field. The overall look is high-end sports film realism with dramatic atmosphere and flawless visual integration.
Greek god bronze boxing ring:
Create a smooth cinematic camera movement inside a realistic professional boxing ring arena. My character stands in the center of the ring under authentic stadium lighting. The camera performs a slow, steady circular rotation around him at chest height, maintaining consistent distance and smooth stabilized motion. The movement is gentle and controlled, as if mounted on a motorized dolly, with no sudden acceleration and no handheld shake. The character remains fully visible and centered . Lighting is realistic arena white light with natural shadows on the canvas. The rotation is continuous and fluid, emphasizing atmosphere and presence without dramatic camera drops or aggressive movement. The overall feeling is cinematic, controlled, and immersive, focused on realism and smooth rotational motion only.
Keeping product text consistent with Nano Banana and Kling
This tutorial shows how to build product prompts that preserve small branding details and keep typography accurate from the very beginning.
Using Nano Banana Pro inside the Image Generator, the workflow focuses on a running shoe with multiple subtle text elements, making it a strong case study for typography control. The tutorial explains how prompt structure matters more than prompt length, organizing information in a clear hierarchy: composition, exact product identity, typography control, true color, lighting harmony, optical clarity, and negative constraints.
Two still life examples demonstrate how the same structure can maintain readable, stable branding across different compositions and lighting setups. The process then moves into the Video Generator with Kling 3.0 Omni, showing how multiple image references can reinforce product identity and typography even when motion and perspective changes are introduced.
By the end, the tutorial makes clear that consistent product typography is not a matter of luck, but the result of structured prompting and strong reference control.
The products and brand names shown in this tutorial are used for illustrative and educational purposes only. All trademarks, logos and brand names are the property of their respective owners. Their inclusion does not imply any affiliation, endorsement or sponsorship by or with Magnific.
Prompts
First shoe:
Create a premium editorial product photograph inspired by a luxury fashion still life. Match the reference composition exactly: one single shoe placed diagonally and balanced naturally on a small stack of rough sandstone rocks, minimal scene with a smooth gradient sky background, warm sunlit highlights and clean sculptural shadows. The shoe must be the Adidas running shoe from the provided references, with exact proportions, silhouette, mesh texture, panel edges, stitching, heel shape, laces, and outsole profile. Preserve true color accuracy: deep cobalt blue upper, turquoise heel collar and rear panel, neon lime laces and neon lime toe edge accent, off white midsole, thin mint green band along the upper midsole edge, black outsole sections, three light gray stripes on the lateral side. Typography must be perfectly readable, sharp, and placed exactly like the reference shoe: on the lateral heel side panel, small light gray text reading ADIZERO TAKUMI SEN 10 aligned along the diagonal heel overlay; on the lateral midsole near midfoot, small light gray text reading LIGHTSTRIKE PRO horizontally on the foam; no misspellings, no extra letters, no warped glyphs, no duplicated text, same font weight and spacing as the reference. Lighting: soft directional daylight from upper left, gentle falloff, crisp but not harsh, with subtle warm highlights on the rocks and slightly cooler shadows to harmonize with the shoe’s blue and lime accents. Color palette should feel controlled and harmonious: earthy warm rocks, soft beige highlights, cool blue undertones in shadow, allowing the neon lime accents to pop without clashing. Shallow depth of field, tack sharp focus on the shoe and its typography, background softly blurred, premium fashion product photography mood, clean contrast, refined tonal balance, no oversaturation, no heavy color grading. Negative: avoid any other text or logos, avoid incorrect spelling, avoid melted typography, avoid extra stripes, avoid plastic looking rocks, avoid glossy wet surfaces, avoid harsh specular hotspots, avoid noise, avoid motion blur.
Second shoe:
Create a premium editorial product photograph in the exact composition style of the reference still life: one single shoe placed diagonally and leaning naturally against a stack of large terracotta clay forms, curved pipes and sculptural cylinders, warm matte clay with subtle porous texture and small imperfections, earthy orange-brown tones. The product must be exactly my running shoe Adidas Adizero Takumi Sen 10 matching the provided white background references with perfect proportions, mesh upper texture, panel edges, stitching, lace routing, sole shape, heel and collar shape, and the three Adidas stripes placement and angle. Typography must be perfectly accurate, crisp, and fully readable with exact placement: on the lateral rear side overlay the text ADIZERO TAKUMI SEN 10 in small uppercase light gray, aligned along the diagonal heel overlay; on the lateral midsole near the midfoot the text LIGHTSTRIKE PRO in small uppercase light gray; on the tongue label the text ADIZERO printed vertically in bold uppercase black; inside the insole the large text ADIZ in bold uppercase black; on the outsole lime section the text CONTINENTAL in small black with the small logo mark beside it. Preserve true color accuracy: deep royal blue upper, bright neon yellow laces, turquoise mint collar and heel lining, off white midsole, black outsole areas, lime outsole forefoot section, and the three stripes in light silver-gray. Lighting must be soft directional daylight from one side, gentle sculptural shadows and warm highlights across the clay while keeping the shoe crisp and clean, with slightly cooler undertones in the clay shadows to complement the shoe blue and lime. Color palette must feel analog but controlled: warm terracotta highlights plus soft beige neutrals plus cool blue and lime accents, no harsh grading, no oversaturation. Shallow depth of field, tack sharp focus on the shoe and its typography, clay background softly blurred. High-end luxury minimal editorial mood, clean contrast, refined tonal balance. Negative constraints: do not change the shoe design, do not invent extra logos or text, do not misspell any typography, no floating shoe, no additional shoes, no extra props, no reflections that distort letters, no watermark, no frame, no heavy vignette, no artificial CGI look.
Shoe short prompt:
Create a premium editorial still life inspired by the reference terracotta composition. My Adidas Adizero Takumi Sen 10 placed diagonally and leaning naturally against warm matte terracotta clay forms. Use the exact shoe design from the provided references, preserving proportions, materials, and stitching. All typography must be perfectly accurate and readable in its correct position: ADIZERO TAKUMI SEN 10 on the heel overlay, LIGHTSTRIKE PRO on the midsole, ADIZERO on the tongue, ADIZ inside the insole, CONTINENTAL on the outsole. Maintain true color accuracy of the royal blue upper, neon yellow laces, mint collar, off white midsole, and lime outsole. Soft directional daylight, warm highlights, slightly cool shadows for balance. Shallow depth of field, sharp focus on the shoe. No design changes, no misspellings, no extra logos.
Discover Kling 3.0 Omni
This tutorial explores Kling 3.0 Omni, the most advanced version of the Kling video model.
You’ll learn how Omni enhances cinematic realism with stronger character consistency, improved lip-sync, and more accurate prompt interpretation. Through three practical examples, the tutorial demonstrates how to combine image references, dialogue, object replacement, and video references to achieve stable, high-quality results.
First, you create a dialogue scene using multiple image references and native audio generation for expressive lip-sync. Then, you modify a specific object within the scene using reference-driven transitions while maintaining character identity. Finally, you replicate complex motion using a video reference, combining movement accuracy with cinematic continuity.
By the end, you’ll understand how Kling 3.0 Omni extends the capabilities of the base model and how to leverage references and audio to achieve precise, controlled video generation.
Prompts
Example 1- Lip sync
1 - A wide cinematic establishing shot of a vast Texas ranch like @Ranch at golden hour, with dusty open fields, long wooden fences, and dramatic warm sunset light casting long shadows. @Johny The Boy stands beside his horse while @Little Timmy sits on the ranch fence. The camera begins as a wide shot and slowly performs a controlled zoom-in toward the two characters, building intimacy and tension.
Sound: Natural ranch ambience, soft wind, distant wooden creaks, subtle horse movement, and low indistinct murmuring between them. (4s)
2 - A dramatic cinematic close-up of @Little Timmy, golden light wrapping around his face while the ranch background remains softly blurred. The camera holds steady to capture micro-expressions and sincerity.
Dialogue (10-year-old boy voice, perfectly lip-synced): “Hey Johny ‘The Boy’, when I grow up, I’d like to be a great cowboy like you…”
Lip movement must precisely match each word with natural child articulation and timing.
Sound: Clear child voice layered over subtle ranch wind. (5s)
3 - A powerful cinematic close-up of @Johny The Boy his weathered face half-lit by warm sunset light, shallow depth of field isolating him from the ranch. The camera remains steady to emphasize authority and presence.
Dialogue (deep, rugged voice, perfectly lip-synced): “Little Timmy, you are an AI-generated character. You can become anything you want with a good prompt…”
Lip synchronization must be precise and natural, matching speech rhythm.
Sound: Deep resonant voice with faint wind and distant ranch ambience. (5s)
Example 2 - Object reference
A cinematic close-up of a rugged Texas cowboy @Johny The Boy standing in a vast, sunlit ranch landscape at golden hour. He is a middle-aged man with weathered skin, short salt-and-pepper beard, intense eyes, and a classic wide-brim cowboy hat. Warm sunset light casts strong shadows across his face while dry wind moves faint dust through the air. Wooden fences and open, arid fields stretch behind him, with no people visible. The camera performs a slow, controlled zoom out until his full face and entire hat are clearly framed, maintaining shallow depth of field and dramatic contrast.
He lifts one hand and calmly touches the brim of his hat in a subtle greeting gesture. The moment his fingers make contact, the hat seamlessly transforms into @Hat1 , keeping perfect lighting continuity and realistic texture. After lowering that hand, he raises the other and touches the brim again. On contact, the hat smoothly transforms into @Hat2, with no distortion and natural perspective alignment. The camera remains steady, emphasizing realism, tension, and western cinematic atmosphere as dust drifts in the warm light.
Example 3 - Video reference
Generate a video that exactly matches the action, camera movement, timing, framing, pacing, and shot composition from @Video. All motion, gestures, performance, and progression must replicate it precisely. Replace the characters, horse, and environment with those from @Johny Horse, preserving their exact appearance, lighting, textures, and proportions. Keep the identical choreography and cinematography. No reinterpretation or additional changes, same structure, new visual assets.
Bring your sketches to life with Nano Banana Pro
This tutorial demonstrates how to use Nano Banana Pro to transform hand-drawn sketches into realistic, production-ready visuals.
Starting from a simple chair sketch, you learn how to guide the model toward a physically plausible, high-end design object by defining materials, lighting, weight, and camera perspective. The tutorial shows how increasing prompt specificity gives tighter control over realism while preserving the original structure of the drawing.
The workflow then expands to an interior sketch, converting abstract architectural lines into a believable photographic space. Finally, both generated elements are combined inside the Image Generator, demonstrating how independently developed assets can integrate naturally while maintaining scale, lighting consistency, and material coherence.
By the end, you understand how to build modular design elements from sketches and assemble them into cohesive scenes using Nano Banana Pro.
Prompts:
Chair:
Transform the provided sketch into a hyper-realistic, premium design object, preserving the exact silhouette, curvature, and proportions of the original sketch while translating it into a physically plausible, high-end manufactured piece. The object should feel sculptural, refined, and architectural, as if designed by a top contemporary furniture studio.
Materiality is premium and tactile: choose a smooth molded composite or solid surface material microcement with subtle surface variation, soft micro-texture, and realistic imperfections. The material should read as heavy, dense, and well-crafted, with rounded edges that catch light naturally and smooth transitions between planes.
The form must exhibit accurate spatial volume and weight, grounded firmly on the surface below. The curvature should feel continuous and intentional, with elegant tension in the twist of the structure. No sharp edges, no visual noise—pure sculptural clarity.
Lighting is studio-quality, architectural lighting: a soft, directional key light from one side to reveal curvature and depth, complemented by gentle fill light to avoid harsh shadows. Subtle contact shadows anchor the object to the ground plane. Highlights are controlled and diffused, never glossy or blown out. The lighting should emphasize form, material quality, and realism.
The object is isolated in space, placed on a neutral, soft background—light warm gray, pale stone, or muted off-white—with a very subtle gradient. The background must enhance contrast and readability without competing with the object. No environment reflections, no context yet.
Camera is set at a three-quarter angle, slightly above seat height, with a realistic focal length (around 50–70mm equivalent) to avoid distortion. Depth of field is moderate, keeping the entire object sharp while maintaining a natural photographic feel.
Overall aesthetic is contemporary, refined, and gallery-level, suitable for high-end interior design visualization. The result should look indistinguishable from a real, professionally photographed design prototype.
No people, no environment, no text, no logos, no stylization. Pure object realism, premium materials, perfect lighting.
Silla real + Interior real juntos:
Place the chair @img2 inside this room @img1 , paying attention to proportion and composition. place it facing the camera in the top left corner
Prompt interior + Sketch reference
Transform the hand-drawn interior sketch into a fully realistic, high-end living room, strictly preserving proportions, layout, and architectural intent. The result must read as a real, inhabitable space designed by a contemporary interior studio, never as a render.
The space is calm, sculptural, and refined, built with premium, tactile materials: warm mineral plaster or limewash walls with subtle texture; light natural stone, polished microcement, or pale oak flooring with visible grain and imperfections; built-in elements and feature walls in warm wood tones with fine vertical grain, matte finish, and precise joinery; dense, soft textiles such as woven rugs, upholstered volumes, and heavy curtains with natural drape.
The color palette is warm, neutral, and cohesive—creams, sand, stone, warm greige, muted beige, light clay, and gentle wood browns. No saturated colors or strong contrasts; the atmosphere is quiet, timeless, and harmonious.
Lighting is purely natural and architectural. Soft daylight enters through large floor-to-ceiling windows, producing realistic falloff, gentle shadows, subtle ambient bounce, and controlled highlights. No artificial lighting is used; contact shadows ground all elements naturally.
Scale and weight are accurate. Furniture feels heavy, grounded, and ergonomic, with smooth curvature and intentional negative space. Nothing floats or feels decorative without purpose.
The camera is at human eye level with a realistic focal length (35–50mm), clean undistorted perspective, and natural depth of field, similar to professional interior photography.
The final image must be indistinguishable from a real photograph, with correct material response, subtle imperfections, realistic light behavior, and quiet realism.
Kling 3.0 multi-shot mode
This tutorial explores how to use Kling 3.0 in multi-shot mode to create structured, cinematic sequences within a single generation.
You learn how to build progression across multiple shots, defining each scene independently while maintaining visual continuity. The tutorial demonstrates how to structure prompts for environment, character, action, camera movement, cinematic detail, and ambient sound, all within the 512-character limit per shot.
It also explains how to assign durations to individual shots while keeping the total sequence under 15 seconds, allowing you to shape rhythm and pacing like an editing timeline. A second example introduces a Start Image, showing how multi-shot mode preserves composition and lighting while focusing prompts on action and camera behavior.
By the end, you understand how to construct cohesive multi-scene videos with precise control over movement, transitions, and narrative flow using Kling 3.0.
Prompts
Rugby Player
1- A fast, aggressive cinematic close-up tracking shot of an All Blacks player with Māori features, long dreadlocks, traditional tattoos, and two black face paint lines. Hollywood-style night rugby match under intense stadium floodlights. The camera moves dynamically alongside him to amplify speed and adrenaline. The player is sprinting powerfully with the ball, defenders dive unsuccessfully. Sweat, turf stains, and motion blur enhance realism while the roaring crowd and pounding boots dominate the sound. (3s)
2- Slow-motion low-angle close-up of his explosive footwork cutting across the grass. His cleats dig into the turf, sending blades of grass and dirt into the air in crisp detail under bright white lights. The camera stays tight to emphasize agility and control as opponents miss by inches. The scrape of studs on turf is sharp while the crowd noise swells behind. (3s)
3- A wide, high-speed dynamic tracking shot of the player charging toward the goal line. The camera sweeps rapidly to heighten velocity and tension as defenders chase desperately behind without closing the gap. Stadium lights cast dramatic highlights across his tattooed arms and determined expression. The crowd roar intensifies with anticipation. (3s)
4- Slow-motion extreme close-up of the player grounding the ball firmly on the grass to score. The turf compresses beneath the ball as dirt and grass rise on impact. His tattooed forearm tightens with effort under bright floodlights. The stadium erupts in celebration, triumphant cheers echoing through the night. (3s)
Coffee
1- A slow-motion cinematic close-up of the barista unscrewing the portafilter handle from the espresso machine. Warm sunrise backlight highlights drifting steam and floating particles in the air. The camera remains steady, capturing the subtle rotational movement and mechanical precision with detailed texture and cinematic realism. Sound: Metallic twisting of the portafilter, low espresso machine hum, soft café ambience in the background (3s)
2- A slow-motion macro close-up of the portafilter basket filled with ground coffee. The barista presses it firmly with a tamper in a controlled motion. Fine coffee grounds compress under pressure, illuminated by warm golden light. Sound: Soft scrape of metal on coffee grounds, dense pressing sound, distant café murmur. (3s)
3- An slow-motion close-up of espresso extraction. Dark coffee drips and flows smoothly into a clean ceramic cup. Steam curls upward in the warm backlight while the camera holds steady to emphasize texture. Sound: Rhythmic dripping of espresso, steady mechanical vibration of the machine, subtle café background noise. (3s)
4- A slow-motion close-up of a metal pitcher pouring milk into the cup. The camera remains tight and stable, highlighting the smooth flow of milk, subtle microfoam texture, and condensation on the metal surface in cinematic detail. Sound: Gentle hiss of steam, soft swirling of milk mixing with espresso, calm café ambience. (3s)
5- A cinematic slow-motion macro close-up of the finished coffee cup filled to the rim with textured milk foam. Steam rises gently as golden light enhances the glossy surface and creamy texture, emphasizing quality and craftsmanship. Sound: Soft ceramic contact as the cup is placed on the counter, low ambient café atmosphere. (3s)
Kling 3.0 single-shot mode
This tutorial walks through how to use Kling 3.0 in single-shot mode inside the Video Generator. The process begins with a prompt-to-video example, showing how detailed and intentional writing directly affects motion, framing, and cinematic quality when no Start Image is used.
You follow a practical prompt structure that covers shot type, character, environment, action, camera movement, and final visual detail. The tutorial also explains how Kling responds to film language such as dolly-in or handheld tracking to achieve more controlled and realistic movement.
A second example introduces a Start Image, shifting the focus from visual description to action and camera direction. You also review key generation settings including resolution, duration, format, and optional sound.
By the end, the fundamentals of working with Kling 3.0 are clear, providing a solid base for experimenting with cinematic video generation.
Prompts:
Climbing Girl
An extreme close-up of a 30-year-old red-haired woman with tight braids, slim and highly athletic, wearing rolled-up jeans and a sport top. She is halfway up the vertical wall of El Capitan in Yosemite during golden hour sunrise, warm orange light illuminating the granite surface and her skin. Her face shows intense physical effort as she clings to the rock. The camera rapidly tilts and drops to a tight close-up of her chalk-covered hand gripping a small rock edge, then pulls back in a fast drone-like movement revealing her suspended high above the valley with a wide view of the landscape. Fine chalk dust drifts into the glowing morning light, emphasizing the height, tension, and scale of the scene.
Porsche Desert
A car racing through the desert, highly dynamic camera movement to emphasize speed and adrenaline. The camera begins in a wide side-profile tracking shot, moving aggressively parallel as dust explodes behind. Without cutting, it rapidly accelerates forward into a front-facing position and performs a fast zoom-in to the windshield, revealing the driver’s intense expression as the background streaks with motion blur. In one sharp whip movement, the camera drops into an extreme macro close-up of the spinning wheel, capturing sand and stones spraying backward. The camera then surges upward into the sky in a fast crane movement.
Mastering Templates
This episode brings together everything you’ve learned throughout the course through the use of templates.
You’ll explore how templates work across image, video, and audio generation, and how they capture complete setups including prompts, references, and settings. The episode shows how templates help you move faster, stay consistent, and avoid starting from scratch every time.
By the end of this episode, you’ll understand how to use and create templates as reusable starting points, turning individual creations into reliable workflows.
Using reference video to control movement
This episode shows how video references can be used to control movement in video generation.
You’ll learn how an existing motion sequence can guide the way a character moves, independent of how that character looks. The episode focuses on transferring rhythm, weight, and physical behavior from a reference video to a generated figure.
By the end of this episode, you’ll understand how video references help achieve more natural, coherent motion and reduce unpredictability in animated video results.