Карту глубины можно сгенерировать тут
Главная фишка здесь именно в контрасте размеров. Я не уменьшаю одного человека уже внутри готового видео и не растягиваю второго. Сначала отдельно подготавливаю двух персонажей с правильной внешностью, одеждой и оружием, затем использую карту глубины как основу для камеры и композиции, а финальную сцену собираю в Seedance 2.5.
Для работы нужен аккаунт в Syntx.ai, тариф PRO или выше, две фотографии людей, карта глубины и звук, который при необходимости добавляется уже при монтаже. В исходной схеме для видео используется Seedance 2.5
Ниже я полностью сохраняю этапы генерации и исходные промты, потому что именно в них зафиксированы размеры персонажей, положение мечей, направление ветра и движение камеры.
Почему я сначала готовлю персонажей отдельно
Если сразу загрузить две обычные фотографии в видеогенератор и попросить сделать одного человека маленьким рыцарем, а второго огромным персонажем, модели придется одновременно решать слишком много задач.
Ей нужно сохранить два лица, придумать броню, спортивную одежду, два разных меча, правильно изменить относительный масштаб людей, поставить их в нужные позиции и еще воспроизвести сложную погоду.
Поэтому я сначала получаю две чистые карточки персонажей.
Первый герой уже выглядит как небольшой рыцарь. Второй уже подготовлен как крупный человек в сером спортивном костюме с огромным мечом.
После этого Seedance не придумывает их дизайн заново, а использует готовые изображения как визуальную основу.
Шаг 1. Готовим фотографии персонажей и карту глубины
Первого персонажа я создаю через Seedream 5.0 Pro. Исходная фотография отправляется в раздел изображений, формат ставится 3:4, качество - 2K. Именно такая последовательность указана в исходной инструкции.
Промт для первого персонажа
Промт оставляю без изменений:
Create a photorealistic full-body character reference image using as the identity reference.
[IDENTITY]
Preserve the person's recognizable face, gender presentation, actual apparent age, skin tone, hairstyle, hair color, and facial hair where present. Keep their natural facial anatomy. Do not turn them into a child, an animal, a cartoon, or a different person.
[SMALL HUMAN CHARACTER]
Reimagine this person as a miniature human knight for a cinematic fantasy scene. Use a compact, small-stature silhouette with believable human anatomy and restrained stylization. Keep a naturally proportioned head and recognizable adult features if the reference person is an adult. No oversized baby head, chibi proportions, feline ears, paws, fur, or tail.
[ARMOR AND CLOAK]
Replace the original clothing with a coherent suit of weathered steel knight armor: fitted breastplate, modest shoulder plates, articulated forearm protection, protective gloves, a dark fabric underlayer, and armored boots. Add a short, worn charcoal-gray cloak attached at the shoulders and falling behind the body. The cloak has a gently wind-lifted edge toward screen-left.
No helmet or hood covering the head. Keep the face and hairstyle fully visible. No oversized shoulder armor, excessive ornament, glowing runes, or bulky shapes hiding the human figure.
[SWORD AND POSE]
The person stands firmly, facing almost directly toward the camera, with feet slightly apart. Hold one straight steel sword vertically in front of the torso, POINT UP, using both hands around the hilt. The hilt is around the upper abdomen; the blade rises above the head without blocking the eyes, nose, or mouth. Offset it slightly from the facial centerline if needed.
The sword is proportionate to this smaller knight, with a simple crossguard, dark grip, and realistic metal surface. Both hands remain anatomically clear. No extra fingers, floating weapon, duplicated blade, or blade passing through the face.
[EXPRESSION]
Serious, composed, quietly determined. Relaxed closed mouth, no smile, no visible teeth. Gaze forward.
[REFERENCE IMAGE PRESENTATION]
One character only, centered, full body from head to soles. Include the entire sword tip and cloak with clear margins around the silhouette. Neutral gray studio background and ground, soft directional light, subtle contact shadow, readable skin and armor detail.
Use restrained perspective without wide-angle distortion. This is a clean character reference, not an action scene: no fog, grass, storm, motion blur, text, labels, collage, or additional views.
Vertical 3:4 composition. Photorealistic face, believable human anatomy, detailed fabric and metal.
На этом этапе мне важно не получить максимально красивую фэнтези-картинку, а добиться хорошего референса персонажа.
Человек должен быть показан полностью, меч не должен закрывать лицо, а само лицо должно оставаться узнаваемым.
Фон специально остается нейтральным. Туман, поле, дождь и драматический свет появятся уже в видео.
Готовим второго персонажа
Для второго человека исходная схема использует GPT Image 2.5 Sunburst, снова с соотношением сторон 3:4 и качеством 2K.
Здесь мне нужен уже не рыцарь. Второй персонаж должен выглядеть как большой человек в простом сером спортивном костюме.
Важно, что модель не должна физически уродовать его ради ощущения масштаба. Голова, руки и ноги остаются нормальными. Огромным он будет выглядеть только относительно первого героя.
Промт для второго персонажа
Его тоже оставляю без изменений:
Create a photorealistic full-body character reference image using @image1 as the identity reference.
[IDENTITY]
Preserve the person's recognizable face, gender presentation, apparent age, skin tone, hairstyle, hair color, facial hair where present, and distinctive facial anatomy. Keep the person clearly identifiable.
[LARGE HUMAN CHARACTER]
Present this person as the larger, imposing human counterpart to a miniature knight. Use a strong, grounded silhouette while respecting the reference person's natural build. Do not force masculine features onto a woman, exaggerate muscles, enlarge the head, or distort the limbs.
The person should have normal human anatomy and proportions; their greater scale will be established relative to the smaller character in the final scene.
[GRAY OUTFIT]
Replace the original clothing with a coordinated gray casual sports outfit:
a plain medium-gray pullover hoodie,
matching gray jogger trousers,
and gray athletic sneakers.
The hood is DOWN, resting behind the neck and on the upper back. It must not cover the head, hair, ears, or forehead.
Use realistic cotton-fleece texture, natural folds, ribbed cuffs, and a comfortable fit that preserves the person's body shape. Keep the sneakers believable and proportionate.
No armor, cape, coat, helmet, visible brand logos, or added jewelry.
[GREAT SWORD AND POSE]
The person stands firmly, facing almost directly toward the camera, with feet slightly apart. Hold one enormous, broad steel greatsword vertically before the torso, POINT UP, with both hands around its long grip.
The hilt is around the upper abdomen. The blade extends well above the head and is substantially larger and broader than a normal sword.
Give it a coherent straight blade, simple strong crossguard, dark wrapped grip, and believable weathered metal. Keep it large but controllable within the composition.
Do not let the blade cover the face: position it slightly off the facial centerline while keeping it upright before the body.
Show correct hand contact, natural elbows, and a balanced stance appropriate to supporting the weapon. No bent blade, duplicated sword, floating grip, extra fingers, or intersecting anatomy.
[EXPRESSION]
Serious, calm, resolute. Gaze forward. Lips naturally closed, no smile and no visible teeth. No exaggerated aggression.
[REFERENCE IMAGE PRESENTATION]
One character only, centered, full body from head to sneaker soles. Include the complete sword tip with sufficient space above it. Neutral gray studio background and ground, soft directional light, subtle contact shadow, clearly readable face and outfit.
Use restrained perspective without wide-angle distortion. Match the clean presentation of a character reference image.
No fog, grass, storm, motion blur, text, labels, collage, or additional views.
Vertical 3:4 composition. Photorealistic skin, natural human anatomy, detailed gray fabric and steel.
Здесь я особенно смотрю на меч и руки. Если оружие уже на референсе получилось кривым или пальцы плохо держат рукоять, в видео эта ошибка может только усилиться.
Поэтому плохой референс лучше переделать сразу.
Зачем в этом тренде нужна карта глубины
Карта глубины здесь работает не как еще одна картинка персонажа.
Я использую ее для другого: она показывает Seedance, как должна двигаться камера и где должны находиться люди относительно друг друга.
В начале маленький рыцарь занимает заметную часть кадра. Затем камера физически уходит назад, и только после этого слева открывается второй, значительно более крупный персонаж.
За счет этого и появляется главный визуальный эффект тренда.
Если заменить такое движение обычным цифровым уменьшением изображения, сцена будет выглядеть значительно слабее. Зритель должен чувствовать изменение перспективы.
Шаг 2. Создаем видео в Seedance 2.5
Дальше я перехожу в раздел видео.
Исходные настройки оставляю точно такими же:
выбираю Seedance 2.5;
режим Omni Reference;
качество 720p;
при необходимости 480p для экономии или 1080p для максимального качества;
формат 4:3;
длительность 10 секунд;
загружаю первого персонажа, второго персонажа и карту глубины.
Карту глубины можно сгенерировать тут
Достаточно взять мое видео и закинуть в этот сервис
После вставки промта в интерфейсе нужно проверить привязку @image1, @image2 и @video1.
В исходной инструкции указано, что при
ручной работе эти обозначения нужно удалить и добавить заново через @, выбрав соответствующий файл. В Syntx.ai референсы привязываются автоматически.
Как распределяются файлы
Здесь важно ничего не перепутать.
@image1 - маленький рыцарь справа.
@image2 - большой персонаж слева.
@video1 - карта глубины, которая задает камеру, расположение героев и момент раскрытия второго человека.
Именно такое распределение зафиксировано в исходном видеопромте.
Промт для Seedance 2.5
Сам видеопромт я тоже оставляю полностью без изменений.
[TASK — TWO CHARACTERS IN A VIOLENT WINDSTORM]
Create a 10-second cinematic photorealistic video, horizontal 4:3, 24 fps output.
One continuous handheld shot in a violently windswept, fog-covered grassland.
Begin on the first character holding a sword vertically before their body. Move the camera backward to reveal the second, larger character on the left holding a much larger sword upright.
Both remain serious and resolute while fast-moving fog, fine wind-driven rain, grass, and loose fabric visibly stream from SCREEN-RIGHT TO SCREEN-LEFT.
[REFERENCE ASSIGNMENTS]
@video1 is a DEPTH-MAP reference for camera trajectory, framing progression, spatial arrangement, approximate poses, sword placement, and reveal timing.
@image1 defines the complete appearance of the FIRST, SMALLER character, initially featured and positioned on the RIGHT in the final composition.
@image2 defines the complete appearance of the SECOND, LARGER character, revealed on the LEFT.
Each photograph defines gender presentation, apparent age, face, skin tone, hairstyle, body proportions, clothing, accessories, and footwear.
Preserve any miniature or larger-scale character design already established in the prepared images.
[DEPTH — SOFT SPATIAL GUIDANCE]
Use @video1 as a moving spatial storyboard, not a rigid mesh to texture.
Follow the backward camera movement, changing viewpoint, relative positions, and left character's reveal.
Do not inherit the source figures' facial anatomy, headwear, armor silhouettes, or clothing when these differ from the photographs.
Photos take priority for character appearance. The written weather instructions take priority for fog, rain, and wind intensity.
Do not copy static or weakly moving fog from the depth map. Generate the clearly visible, fast-moving weather described below.
Adapt the source composition to 4:3 without stretching the image or cropping out the left character, faces, or important sword details.
[IDENTITY AND CLOTHING]
Construct each complete character from their corresponding photograph.
Preserve the face, physique, relative character design, hairstyle, outfit, and accessories throughout.
Do not paste replacement faces onto the source bodies.
Do not blend or swap identities.
Preserve armor and a cloak if present in @image1.
Preserve the gray hoodie, trousers, sneakers, and lowered hood if present in @image2.
Do not raise a lowered hood or invent additional clothing.
Wind animates the existing garments without changing their construction.
[PLACEMENT AND REVEAL]
The first character is visible at the opening and remains on the RIGHT as the shot widens.
The second character is already standing to the LEFT, outside the initial framing, and is revealed by the camera's backward movement.
Preserve the intended contrast between the smaller right character and larger left character without distorting their individual anatomy.
Neither character walks into position, materializes, fades in, or changes size during the shot.
[SWORDS]
Both hold swords vertically before their bodies with the POINTS UP and both hands around the hilts.
The right character holds the smaller sword. The left character holds the much larger, broader greatsword.
Preserve sword designs from the prepared images where visible.
Keep hands near the upper abdomen or lower chest, following the reference pose.
Blades remain rigid and coherent. Keep faces readable rather than hidden behind the blades.
No swinging, attacking, bending weapons, duplicated swords, or fingers intersecting the grips.
[PERFORMANCE]
Serious, calm, determined expressions. No smile, dialogue, or exaggerated aggression.
Preserve reference head angles, attention, and sword-holding posture.
Allow subtle breathing, blinking, and restrained balance corrections against the gusts.
The characters hold their ground. No sliding feet, exaggerated swaying, or weightless bodies.
[CAMERA — BACKWARD REVEAL AND STRONG HANDHELD SHAKE]
Follow the backward reveal in @video1.
Begin with the first character and upright sword, then physically move backward with the source-guided reframing to reveal the second character on the LEFT.
End in a wider two-character composition within the 4:3 frame.
Preserve changing perspective and depth relationships. This is not merely a digital zoom-out.
Superimpose pronounced, irregular handheld movement, as if the operator is struggling against severe gusts: brief lateral and vertical jolts, small roll disturbances, and imperfect corrections.
Do not smooth the shot into a stabilized gimbal move.
Keep the subjects readable. No repetitive vibration loop, excessive whip movement, or body deformation.
The camera shake and weather movement are separate effects: shaking the frame must NOT replace actual fog movement through the scene.
One continuous shot. No cuts, orbit, drone ascent, or reverse angle.
[TIMING]
Fit the depth reference's reveal naturally into 10 seconds.
Opening: establish the first character, upright sword, and immediately visible storm movement.
Development: continue backward and progressively reveal the larger character on the LEFT.
Ending: both characters and swords remain readable while strong wind, rapidly streaming fog, fine rain, and handheld movement continue.
Do not freeze the final composition or weaken the storm after the reveal.
[ENVIRONMENT]
An exposed, uneven grassland with low vegetation, patches of damp earth, and a distant horizon obscured by moving fog.
Muted gray-green grass, dark earth, pale gray mist, and an overcast sky.
No buildings, extra people, armies, vehicles, dramatic ruins, or invented landmarks.
[CRITICAL WEATHER — VISIBLE HIGH-SPEED RIGHT-TO-LEFT FLOW]
The dominant wind blows continuously from SCREEN-RIGHT TO SCREEN-LEFT.
It must be unmistakably visible in the movement of fog, fine rain, grass, hair, and fabric, not merely implied by sound or camera shake.
The viewer must be able to follow individual fog streaks and denser wisps entering from the right, racing behind or around the characters, and exiting to the left.
This lateral motion must remain visible relative to the characters and terrain even during handheld camera movement.
No stationary background haze, slowly breathing smoke, or fog that only changes opacity in place.
[FOG — FAST STREAMS AT MULTIPLE DEPTHS]
Build the fog from distinct moving layers rather than one uniform gray wall.
BACKGROUND: elongated pale-gray fog banks sweep rapidly RIGHT TO LEFT behind the characters. Visible variations in density reveal their travel across the landscape.
MIDGROUND: more defined wisps and ribbons race between terrain features, behind shoulders, and around the figures, curling and stretching under gusts.
GROUND LEVEL: low fog surges horizontally through the grass and around legs.
OCCASIONAL FOREGROUND: thin translucent strands cross the lower frame quickly, adding depth without hiding faces.
During stronger gusts, distinct fog structures cross a substantial portion of the frame within roughly one to two seconds. Keep this readable as lateral advection, not rapid dissolving or flickering.
Use turbulence and local curls within the dominant RIGHT-TO-LEFT flow. Do not reverse the overall direction.
Provide enough tonal variation behind the moving mist to make its motion visible, while keeping the distant setting obscured.
Do not simulate this by sliding a flat background plate. Fog flows through the scene while the landscape retains coherent geometry.
[FINE WIND-DRIVEN RAIN]
Add fine, sparse-to-moderate wind-driven drizzle as a secondary direction cue.
Small droplets form short, subtle streaks travelling predominantly from SCREEN-RIGHT TO SCREEN-LEFT with a slight downward diagonal.
The rain is pushed almost sideways by the wind, not falling vertically.
Use variation in streak size and focus according to depth; avoid identical repeated lines.
Keep it fine rain, not thick cinematic rain ropes, snow, hail, sparks, or flying debris.
Rain must not obscure the faces or overpower the streaming fog.
Avoid large water droplets on the lens that conceal the scene.
[GRASS, HAIR, AND FABRIC]
Grass bends strongly toward SCREEN-LEFT, with visible waves of pressure travelling through it.
Rooted vegetation stays attached and partly rebounds between stronger gusts.
Loose hair and garment edges stream left in the same wind.
Any cloak pulls outward and snaps left with believable weight, folds, and tension at its shoulder attachments.
Heavy garments respond less freely than thin fabric.
Do not make every strand or blade of grass oscillate in perfect synchronization.
Do not replace directional motion with random fluttering.
[WEATHER COHERENCE AND VISIBILITY]
Fog, drizzle, grass, hair, and fabric share the same prevailing RIGHT-TO-LEFT wind.
Gust intensity varies, but the directional flow never disappears.
Keep faces, hands, and swords readable through gaps and translucent layers.
The background fog must visibly MOVE, not simply make the background gray.
No explosion-like smoke bursts, vertical smoke columns, magical particles, lightning, or opaque whiteout.
[LIGHTING AND COLOR]
Cold, diffuse overcast daylight with soft shadows and restrained contrast.
Desaturated gray-green surroundings and believable skin tones.
Subtle highlights reveal moving rain and fog edges without theatrical backlighting.
Preserve photo-defined clothing colors. Maintain readable texture in gray clothing against the gray atmosphere.
Natural metal reflections on swords and any armor.
No warm sunset, artificial glowing eyes, excessive blue skin, or dramatic spotlights.
[IMAGE QUALITY AND CONTINUITY]
Maintain stable faces, bodies, outfits, sword geometry, and hand anatomy through camera shake and weather.
Use natural motion blur for fast-moving mist, drizzle, fabric, and camera jolts.
Do not smear all fog into a featureless wash; retain enough streak and density structure to show its direction and speed.
No duplicated limbs, changing costumes, sliding feet, warped terrain, or rubber-like deformation.
[AUDIO]
If sound generation is supported, use sustained strong wind with irregular gusts, fabric flapping, vegetation rustle, and restrained fine-rain texture.
No dialogue, narration, singing, added music, or exaggerated cinematic impacts.
Weather must remain visually convincing even with the sound muted.
[FINAL PRIORITIES]
Exactly 10 seconds, horizontal 4:3, 24 fps output.
@image1 is the initially featured smaller RIGHT character; @image2 is the larger LEFT character.
Photos define identity, age, gender presentation, anatomy, clothing, and character design.
Depth guides camera, staging, poses, and reveal timing.
One continuous backward reveal with pronounced irregular handheld shake.
Both swords remain upright, points UP.
Clearly trackable, fast-moving fog streams travel RIGHT TO LEFT throughout the shot, especially BEHIND the characters.
Fine nearly horizontal drizzle reinforces the same wind direction.
Grass and loose fabric bend and stream left under severe gusts.
No static fog backdrop, generic haze overlay, or camera shake substituted for weather motion.
No added scenes, captions, subtitles, or watermarks.
Что именно создает эффект тренда
Когда я смотрю на эту сцену, главный эффект появляется не от брони и даже не от гигантского меча.
Он появляется в момент, когда камера отъезжает.
Сначала зритель воспринимает маленького рыцаря почти как обычного человека. Его настоящий размер становится понятен только тогда, когда слева появляется второй герой.
Поэтому я особенно внимательно слежу за тремя вещами: камера должна физически двигаться назад, второй персонаж не должен «вырастать» внутри кадра, а маленький герой не должен уменьшаться во время ролика.
Оба стоят на своих местах с самого начала. Просто второго человека сначала не видно.
Именно это отдельно зафиксировано в промте.
Почему в видео столько внимания уделено ветру
Второй важный элемент - ощущение настоящего шторма.
Просто добавить серый туман на фоне недостаточно.
По задумке ветер постоянно идет справа налево. Это направление должно быть заметно одновременно по туману, мелкому дождю, траве, волосам и плащу рыцаря.
За счет этого все элементы сцены начинают ощущаться частью одного пространства.
Если туман движется влево, а плащ вдруг развивается вправо, эффект быстро ломается.
Именно поэтому большая часть исходного промта посвящена не внешности персонажей, а физике окружающей среды.
Что делать, если вместо отъезда камеры получается zoom-out
Это одна из самых заметных ошибок.
При zoom-out изображение просто визуально уменьшается. Перспектива почти не меняется.
Здесь же камера должна реально двигаться назад.
Передний герой постепенно становится меньше в кадре, меняются отношения между объектами и открывается пространство слева.
В промте это отдельно сформулировано как:
This is not merely a digital zoom-out.
Если первая генерация просто уменьшает весь кадр, я бы в первую очередь усиливал именно блок CAMERA, а не переписывал описание персонажей.
Если большой персонаж появляется из воздуха
По сценарию он уже стоит слева с самого начала.
Просто в стартовом кадре он находится за пределами композиции.
По мере отъезда камеры второй герой постепенно открывается.
Если модель заставляет его войти в сцену, материализоваться или увеличиваться на глазах, это уже другая постановка.
В таком случае особенно важен блок:
Neither character walks into position, materializes, fades in, or changes size during the shot.
Если лица начинают меняться
Здесь я возвращаюсь не к видео, а к исходным карточкам персонажей.
Они должны быть максимально чистыми и понятными.
Для обоих людей желательно:
хорошо читаемое лицо, полный рост, понятная одежда, правильные руки и уже готовое оружие.
Если на исходном изображении лицо маленькое или меч перекрывает половину головы, Seedance получает плохой референс.
Поэтому иногда правильнее переделать первую картинку, чем десять раз запускать одно и то же видео.
Если туман выглядит как статичный серый фон
В промте специально запрещен неподвижный haze.
Туман должен быть разбит на несколько слоев.
На заднем плане идут крупные полосы, в среднем плане более заметные потоки проходят возле персонажей, а небольшие прозрачные участки иногда пересекают передний план.
Главное - зритель должен видеть, что туман физически перемещается справа налево, а не просто висит в воздухе.
Добавляем звук
В исходной схеме звук можно добавить отдельно на монтаже. Для сцены лучше всего подходят сильный порывистый ветер, шорох одежды, движение травы и очень мелкий дождь.
Музыка здесь не обязательна.
Сам шум шторма уже хорошо поддерживает масштаб сцены.
Если модель умеет генерировать звук вместе с видео, соответствующая инструкция уже находится в основном промте.
Итог
Где встречается этот тренд
Такой ролик сейчас используют не только как отдельную фэнтези-сцену с маленьким рыцарем. Этот визуальный прием встречается и в мемных видео с подписями вроде:
«Когда девушка сказала: “Давай я тебе помогу”»
Смысл строится на контрасте. Сначала зритель видит маленького героя, который выглядит самостоятельным и серьезным, а затем камера отъезжает и рядом появляется огромный второй персонаж. За счет разницы в масштабе обычная ситуация превращается в визуальную шутку.
Поэтому один и тот же шаблон можно адаптировать под разные короткие тренды:
«Когда девушка сказала: “Давай я тебе помогу”»
«Я: справлюсь сам. Она через минуту:»
«Когда попросил совсем немного помощи»
«Я и мой друг, который сказал, что просто постоит рядом»
«Когда решил взять поддержку с собой»
В генерации при этом почти ничего менять не нужно. Маленький персонаж остается справа, большой появляется слева после отъезда камеры, а подпись уже задает смысл ролика.
В этом тренде я не пытаюсь заставить видеонейросеть сделать всю работу за один запрос.
Сначала отдельно готовлю маленького рыцаря и большого персонажа. Затем использую карту глубины для постановки камеры и только после этого собираю сцену в Seedance 2.5.
Так модель получает четкое разделение задач:
фотографии определяют персонажей;
карта глубины задает пространство и движение камеры;
промт отвечает за погоду, поведение героев и финальную композицию.
Именно поэтому результат выглядит значительно стабильнее, чем генерация двух случайных людей сразу внутри сложного штормового кадра.
Главное в этом тренде не просто уменьшить одного человека. Нужно создать момент, когда зритель сам понимает масштаб сцены во время отъезда камеры. Тогда маленький рыцарь действительно воспринимается маленьким, а второй персонаж становится визуально огромным без деформации его тела.
Реклама. SYNTX INTELLIGENCE – FZCO. ИНН:105317427000001