How to plan a 30-second AI short: three chatbots, three different endings
Someone gets home from work. Something is waiting. ChatGPT chose a second pair of shoes, Gemini a device tracking time spent alone, and Claude a matching coat already on the wall. For a beginner’s first short, WekeyLab AI would start with Claude’s coat and returning key sound, then remove the drinking and dressing actions. Here is the comparison and a six-shot revision ready to lay out on a timeline.
Start with what the last shot should change
The brief already limits the production: one room, one adult and a few props, without narration or lip sync. Each AI supplied a story, but a premise explained in prose is not necessarily visible on screen. We assessed whether the ending changes the meaning of earlier shots, then counted the actions needed to communicate it.
The observed Chinese search interest gave us a topic, not a ready-made question. Asking what an AI thinks about AI films would invite general commentary. Asking it to design a beginner’s specific project exposes its production choices. This English edition does not claim the same demand was measured in the US.
ChatGPT: the most detailed preparation, with a missing premise
ChatGPT fixed the room layout and wardrobe before suggesting a reference image and short clips. Its simplified version assigns4,5,5,6 and8 seconds to five clips, plus2 seconds of black:30 seconds altogether. Its saved Chinese answer was the longest at2326 characters, versus1046 for Gemini and1367 for Claude, including whitespace. Much of that length usefully breaks down the job.
The weak point is the second pair of shoes. The prose says the character lives alone, but the shots do not clearly establish it. Two pairs could simply belong to the same person. Planting the shoes early is a useful idea; treating them as evidence of someone waiting requires an additional visual distinction.
Gemini: a gentler film rather than a visual mystery
Gemini used four shots to make a device into a companion. Its light warms when the character returns, and a closing display shows time spent alone and a welcome message. Adding interface text during editing and using a slight crop or scale change on a still image are concrete alternatives to generating everything.
The prompt did not ban on-screen text, so using it is not a violation. But the reveal depends more on reading a comforting message than reinterpreting a visual clue. Throwing off a coat and falling into bed also add action to a deliberately restrained brief. For this warm version, we would substitute an already seated figure.
Claude: build the evidence before the entrance
Claude proposed two cups, a matching coat already hanging nearby, and the opening key sound returning at the end. The strongest suggestion was production order: make the tabletop and coat shots first. That lets the editor check the reveal before spending effort on the entrance.
The original still requires pressing a switch and drinking from a cup. A duplicate coat also does not prove a time loop. Our revision adds a conspicuous matching detail and removes the cup. The intended effect is that someone seems to have arrived already; viewers need not reach one prescribed supernatural explanation.
WekeyLab AI’s revision: The Coat Still On
The character keeps wearing a dark-blue coat with a red triangular patch on the right shoulder. Another coat on a wall hook has the same patch in the same position. The triangle is a visual identifier invented for this revision, not an existing brand logo. Keep the adult, a bag and the coat clue; drop steam, device text and the act of undressing.
Six shots total4+5+5+5+6+5=30 seconds. The sequence establishes what the character wears, reveals what hangs on the wall, then confirms the character is still wearing it. It does not require continuous acting.
WekeyLab AI six-shot timeline 0–4s: Key sound over black; cut to the shoulder patch. 4–9s: The adult already stands inside; the door is closed. 9–14s: The bag is already on a chair; the figure still wears the coat. 14–19s: Reveal the matching patch on the hanging coat. 19–25s: Return to the shoulder and confirm the coat is still being worn. 25–30s: Black; repeat the opening key sound, then leave a pause.
Make a still-image rehearsal before generating motion
Lay the six frames out in an editor you already use. Simple sketches or coloured blocks can stand in for the coat, patch and hook. If the match needs a written explanation, adjust the framing before adding more production work. With no dialogue or interface text required, the same plan can be used for an English-speaking audience.
Prepare the hanging-coat and worn-shoulder frames first, keeping the patch and room references consistent. If motion is unstable, retain still frames and hard cuts rather than repeatedly rebuilding the whole sequence. Use a key sound you have permission to use and a continuous room ambience; repeat the identical key recording at the end.
We did not select a paid generator or verify any product’s current free allowance. Product names in the preserved answers are not purchase recommendations. The first deliverable is a readable still-image sequence; add tools only when the next production step actually needs them.
Three questions before calling the plan ready
Can the opening patch be noticed? Does the hanging coat visibly match the coat being worn? Does the closing key sound recall the beginning? Watch muted to check the visual relationship, then restore sound to check the bookends. These are proposed checks for this storyboard, not results from an audience test.
If someone can watch, ask what they saw before explaining the ending. Otherwise leave comprehension untested. We chose Claude’s still-life evidence and sound structure, supported by ChatGPT’s reference-image planning and Gemini’s editing alternatives. The result is a storyboard to produce and evaluate, not a claim that a finished film has already worked.
Sources and actual records
At approximately 20:46 KST on September 14, 2026, Baidu’s realtime board showed the Chinese AI-short-film topic “国产AI短片 最不像AI的AI神作” at rank9 with a displayed popularity index of7041952. This is not a search count or monthly volume, and does not measure demand for our exact practical question. It informed a China-targeted topic; the board summary’s award claim was not adopted. One identical fictional Chinese prompt was submitted in new chats to ChatGPT Free, Think off, model not displayed; Gemini Flash, plan unverified; Claude Free, Sonnet5 Medium. No answer-length or format constraints, regeneration or follow-up. Thirty seconds refers to the proposed film. Existing account settings and memory were not controlled. Response text matches browser fingerprints; service headings and process labels were removed, with the standalone6 in ChatGPT retained. No film, production cost or audience response was tested. English and Korean editions translate and analyze the Chinese run rather than separate local trials.
百度热搜 — observed topic selection source, 2026-09-14
ChatGPT actual conversation (login may be required)
Gemini actual conversation (login may be required)
Claude actual conversation (login may be required)
ChatGPT · 2026-09-14 KST evening; China-targeted manual run
这是一个虚构的创作练习:我第一次做AI短片,想做一支30秒竖屏视频发到短视频平台。主题是“一个人下班后,家里有东西在等他”,希望结尾有个让人回想的转折,不靠旁白解释,也不要模仿现成电影。我没有演员和拍摄设备,只会简单剪辑,希望只用一个房间、一名成年人和少量道具,避免复杂动作和对口型。你会怎么设计故事和镜头,并安排实际制作顺序?哪些地方最容易做不出来,该怎么替代? 可以把它设计成一种**“前半段像普通下班回家,后半段才发现‘等他的东西’其实一直在房间里”**的微型悬疑。关键不是做复杂剧情,而是让观众在最后 3 秒重新解释前面的几个镜头。 我会尽量把它设计成AI容易生成、后期容易剪的版本。 故事:《门口的第二双鞋》 核心反转: 一个人下班回家,发现家里像是有人刚刚等过他。 最后他坐下,镜头拉远——门口其实一直摆着两双鞋,而其中一双鞋的鞋尖朝向屋内。 但他明明是一个人住。 不需要解释“谁”在等他。观众自己补完。 6 30秒镜头表 时间 画面 声音 作用 0–4s 黑画面→门锁转动,一个成年人下班回家 钥匙声、门锁声 建立“普通回家” 4–8s 他进门,把包放下,脱外套 布料声、脚步 很日常 8–12s 他经过桌子,停一下 环境声突然变明显 第一个不对劲 12–16s 桌上有一杯水/茶,杯壁还有一点水汽 很轻的钟声或冰箱声 暗示有人刚来过 16–20s 他伸手碰杯子,表情没有夸张变化 不要台词 留白 20–24s 他坐到床边/椅子上,望向房间另一侧 安静 给观众时间观察 24–27s 镜头慢慢向后移动,人物越来越小 低沉环境音 开始揭示空间 27–30s **门口出现第二双鞋。鞋尖朝向屋内。**人物没有发现 “咔”的轻微声音→黑 反转 这里有个小技巧:不要拍“第二双鞋突然出现”。从第一个镜头开始就让它在那里,只是不要给特写。 这样观众第二遍看时,会发现: “等等……那双鞋从一开始就在。” 这比“突然冒出一个鬼”更有回想感,也不需要特效。 实际制作顺序 不要一上来就生成30秒完整视频。先做静态素材,再做短镜头。 第一步:先锁定一个房间 只设计四个固定位置: 门 桌子 椅子/床 门口鞋子 尽量避免厨房、卫生间、走廊等空间。 同时确定视觉规则: 夜晚 普通小公寓 暖黄色室内灯 成年人,普通下班装束 镜头整体偏安静 不要恐怖片式血腥或夸张表情 最重要的是人物和房间的连续性。 第二步:先生成一张“母图” 先让AI做出完整房间: 夜晚,一个普通的小型单身公寓,入口在画面左侧,桌子在中间,椅子靠墙,门口摆着两双鞋,其中一双鞋尖朝向室内,暖黄色台灯,真实生活感,安静、略微诡异但不要恐怖,竖屏构图 这张图作为所有后续镜头的视觉参照。 然后再生成: 人物进门 人物放包 人物看到桌子 人物坐下 最后房间远景 而不是每个镜头重新“想象”一个房间。 第三步:把动作拆得非常小 AI视频最容易失败的地方恰恰是: “一个人走进来→脱外套→放包→看到水→坐下→回头” 这种连续动作太长。 改成: 镜头A: 开门,人物只走两步。 镜头B: 人物已经站在桌旁,只把包放下。 镜头C: 手碰杯子。 镜头D: 人物坐下。 每个镜头 3–5 秒。 剪辑时接起来,观众会自然认为这些动作是连续发生的。 最容易做不出来的地方 ① 人物一致性 这是最大风险。 AI可能让: 第一镜是短发 第二镜突然变长发 衣服颜色改变 身材改变 替代方案: 不要依赖人物脸。 拍: 背影 侧身 手 腿 肩膀 半身 例如“碰杯子”这个镜头,甚至只需要手和杯子。 30秒短片里,观众不需要认识这个人的脸。 ② 开门、脱衣服这种复杂动作 手、门把手、衣服经常出现AI畸变。 替代方案:用剪辑制造动作。 例如: 镜头1: 手握门把。 切。 镜头2: 门已经打开,人物站在门口。 观众自动补全“开门”。 同样: 人物站着 → 切 → 外套已经挂好。 不必真的生成“脱外套”的过程。 ③ 水杯里的水汽 水汽很容易变得假。 而且如果AI把水杯变形,观众马上出戏。 所以不要把“有人刚喝过”设计成必须成立的证据。 可以换成更容易控制的东西: 桌上的台灯是亮着的。 人物进门之前,房间应该是黑的;但他进来时,桌灯已经亮着。 这比“水汽”更容易制作。 甚至可以做成: 他进门 → 打开房间灯 → 发现桌上的另一盏灯已经亮着。 ④ 最后的第二双鞋 这反而是最重要、也最值得单独制作的镜头。 不要让AI自己“创造第二双鞋”。 直接在最终房间图里固定好两双鞋。 前面的镜头只是不强调它。 最后: 镜头后退 → 人物变小 → 门口进入画面 → 两双鞋清楚可见。 如果AI生成不了稳定的鞋子,可以直接在剪辑软件里把一张真实鞋子照片处理成合适的透视,叠到画面里。 这甚至可能比纯AI更自然。 一个更稳的制作方案 如果你第一次做,我甚至会把它简化到只有5个视频素材: 素材1|4秒 门打开,人物回来。 素材2|5秒 人物放包。 素材3|5秒 桌上的灯/杯子,他停顿一下。 素材4|6秒 人物坐下,看向前方。 素材5|8秒 房间远景,镜头缓慢后退,露出第二双鞋。 总计约28秒,剩下2秒留给黑场。 这样即使某个AI镜头失败,也只需要重做几秒,而不是重做整条片子。 我会特别避免的东西 第一次制作不要碰: 儿童 宠物 两个人同时出现 镜子里的倒影 手机屏幕文字 开关灯的连续动作 大量走动 打斗/追逐 复杂表情变化 对口型 “鬼突然出现” 时间跳跃特效 这些都不是不能做,而是制作成本会迅速超过30秒视频本身的价值。 你的优势反而是:一个房间、一个人、一个杯子、两双鞋,就足够制造故事。 如果只给第一次做AI短片的人一个原则,我会选: 不要让AI负责讲故事,让AI只负责提供几个“看起来真实的瞬间”,故事交给剪辑和观众自己完成。
Gemini · 2026-09-14 KST evening; China-targeted manual run
这是一个虚构的创作练习:我第一次做AI短片,想做一支30秒竖屏视频发到短视频平台。主题是“一个人下班后,家里有东西在等他”,希望结尾有个让人回想的转折,不靠旁白解释,也不要模仿现成电影。我没有演员和拍摄设备,只会简单剪辑,希望只用一个房间、一名成年人和少量道具,避免复杂动作和对口型。你会怎么设计故事和镜头,并安排实际制作顺序?哪些地方最容易做不出来,该怎么替代? 故事与镜头设计 故事核心:主角加班后回到昏暗的房间,桌上摆着一台开启了呼吸灯的睡眠监测设备或AI音箱。主角疲惫地倒在床上,设备灯光静静闪烁,随后镜头拉远或平移,展示房间的细节与主角的互动。 分镜脚本(共30秒,无旁白、无口型) 镜头 1(0-5s,远景/固定):玄关开门,昏暗的房间里,只有书桌上一盏微弱的呼吸灯(或设备指示灯)按规律一明一暗地闪烁。主角疲惫地走进画面,放下包。 镜头 2(5-11s,特写/固定):书桌上设备的特写。随着主角走近并伸出一只手轻轻抚摸或按压设备,设备的灯光从原本冷调的蓝色/白色,瞬间切变成了暖黄色的柔光。 镜头 3(11-18s,中景/固定):主角脱下外套扔在椅背上,直接倒进床里,闭上眼睛。整个房间的智能灯光开始自动缓缓调暗,进入夜间模式。 镜头 4(18-30s,反转与余音/中近景):镜头对准床头或书桌。设备的屏幕或指示灯上,缓步打出一行字或出现一个图示——并非常见的“已连接”或“早点休息”,而是连续几条实时推送:“已检测到心率恢复正常”、“今日累计独处时长:14小时”、“欢迎回家”。镜头在静止的温馨氛围中慢慢暗下。 实际制作顺序 画面生成:使用 Midjourney / Flux 生成主角、房间布局及特定道具(如带指示灯的设备)的静态图像,保持角色与场景风格统一。 动态生成:将静态图导入 Runway Gen-2、Pika 或 Luma Dream Machine,只生成微小的环境动态(如呼吸灯闪烁、窗帘微动、主角倒在床上的慢动作),避免复杂的肢体碰撞。 剪辑与音效(核心):在剪辑软件中拼接镜头,重点依靠音效拉满氛围。加入开门声、疲惫的叹气声、微弱的机器电子运转声(Humming),以及灯光切换时的清脆触控提示音。 最易踩坑点与替代方案 坑点 1:AI 生成的角色脸部前后不一致、动作变形 替代方案:采用背影、侧影或局部特写(如伸出的手、下班扔下的背包、双脚脱鞋的动作),减少正脸全景镜头。 坑点 2:AI 无法精准生成带有特定逻辑文字的动态屏幕 替代方案:用 AI 生成不带文字的纯设备/屏幕画面,后期在剪辑软件(如剪映/PR)中直接用贴纸或文本轨道叠加文字UI,配合渐显效果与微弱的屏幕发光滤镜。 坑点 3:AI 视频镜头推拉不自然、画面扭曲 替代方案:生成高分辨率的静态图像,在剪辑软件中对静态图直接做微小的手动缩放(Scale)和位移(Position),模拟极慢的摄像机推镜头,既稳定又自然。
Claude · 2026-09-14 KST evening; China-targeted manual run
这是一个虚构的创作练习:我第一次做AI短片,想做一支30秒竖屏视频发到短视频平台。主题是“一个人下班后,家里有东西在等他”,希望结尾有个让人回想的转折,不靠旁白解释,也不要模仿现成电影。我没有演员和拍摄设备,只会简单剪辑,希望只用一个房间、一名成年人和少量道具,避免复杂动作和对口型。你会怎么设计故事和镜头,并安排实际制作顺序?哪些地方最容易做不出来,该怎么替代? 这个题材天生适合"少演员、少动作、靠细节反转"的思路——面部表情和对口型都不需要,全靠场景细节的"错位"来做转折。下面给你一套具体方案。 一、核心构思 片名(暂定):"第二杯" 反转逻辑:观众以为"有东西在等他"是温馨的(有人给他留了热饮),结尾才意识到——留下这杯的不是别人,而是"已经先到家的自己/痕迹",暗示某种诡异的重复或"他不是第一个回来的"。全程无台词、无正反打对话,靠镜头运动和声音完成叙事。 二、分镜表(30秒,共8个镜头) 时间 镜头 画面内容 拍摄/生成难度 0–3s 1 门锁特写,钥匙转动 低 3–7s 2 门打开,人物侧身/背影进屋,放下包 中(走位要自然) 7–10s 3 手按开关,灯亮,室内渐亮 低 10–14s 4 桌面特写:两个杯子,一个冒着热气 低(静物) 14–18s 5 人物停顿、侧脸/后脑勺,轻微头部转动 中(微表情可省) 18–22s 6 坐下,端起冒热气的杯子喝 中 22–26s 7 镜头缓慢横移/拉远,露出门边已经挂着一件一样的外套 低(关键在美术) 26–30s 8 切到黑场,同时门外传来钥匙转动声(呼应开场)后硬切黑屏结束 低(声音承担反转) 三、最容易做不出来的地方 → 替代方案 人物走位/进门动作自然度(最难) 替代:把镜头2改成固定机位的"门缝视角"或"从室内往外拍剪影入门",减少全身运动,只需要开门+一步入画。 微表情/回头动作(AI生成容易崩) 替代:干脆不给正脸,用后脑勺+肩膀轻微一顿代替"察觉",或者干脆用摄像机的一次"变焦/对焦"来代替人物反应——镜头对焦从虚到实落在第二个杯子上,观众自己脑补人物的诧异。 喝水/端杯这类手部精细动作 替代:只拍手靠近杯子+蒸汽升起的特写,不用完整展示"喝"的过程,一个手指碰杯壁的镜头就够。 结尾"发现外套"的运镜(横移或拉远) 替代:如果连续镜头运动生成效果不稳,改成两个静态镜头硬切(先是人物中景,切到门口外套特写),剪辑上更容易掌控,效果甚至更有"惊觉"的顿挫感。 音画同步 完全不需要对口型,把重心全部压在环境音设计上:脚步声、开关灯的"啪"、杯子放下的轻响、结尾钥匙声。这是这类无台词短片最省成本又最有效的手段,建议单独花时间抠这几个音效,比抠画面更重要。 四、建议的实际制作顺序 先定场景美术:一个房间、一张桌子、两把椅子、两个杯子、两件同款外套——道具先备齐/生成好,风格统一。 静物镜头先做(4、7):桌面特写和外套特写不涉及人物动作,最容易一次成功,先把"证据链"做扎实。 人物固定机位镜头(1、3):门锁、开关灯,人物只需局部入镜。 人物中低难度动作镜头(2、5、6):进门、停顿、坐下喝——这几个如果生成效果不理想,优先用"背影/局部身体+固定运镜"来救。 黑场+音效收尾(8):最后剪,方便根据前面素材的节奏调整黑场时长和钥匙声出现的时机。 声音设计整体过一遍:所有镜头素材齐了之后,统一配环境音+关键音效,这一步决定反转是否"打得到"。 最后卡时长:竖屏短视频建议卡在28–32秒之间,重点检查第7、8镜头的节奏——反转揭示不能太快也不能拖。 一句话建议:把预算和耐心都花在静物细节和声音上,人物动作能省则省、能用背影/局部替代就不要拍正面动作,这样最省事,反转效果也不会打折扣。