跳到主要内容
Supermarket
返回能力市场
Claude Skill
programming
MIT

vox-director

Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.

Alisa0808Alisa0808
77/ 100

公开评测 · 综合采用结论

核心能力可用,建议在受控范围内试用

查看评测依据 评测我的项目基于公开项目证据,非安全认证或安装推荐
2.1kstars
317forks
最近更新 1个月前
评测生成时间(北京时间)
本报告引擎
v3.10.0
当前引擎
v3.16.0

本报告与当前引擎使用不同规则;原分数不会自动更新,不同版本的分数不宜直接对比。

重新评测此项目

进入后确认来源与额度,提交才会创建任务。

Evaluation report

综合采用结论

77
B
满分 100
值得试用低风险
决策摘要

核心能力可用,建议在受控范围内试用

75%
中置信度
71
文档
100
安全
64
质量
90
活跃
57
采用
  • 基础评测完成+25/25确定性评分与静态安全扫描已完成
  • README 有效证据+16/256,365 个去重后的有效字符
  • 独立证据来源+4/201 类非重复证据,重复文件不叠加
  • 仓库元数据+10/10已取得仓库状态与采用数据
  • 活跃记录+5/5已取得最近提交时间
  • AI 复核+15/15已完成结构化 AI 证据复核
How it works · 未生成图示

图示结构未通过校验

AI 返回了图示候选,但字段数量、长度或格式未满足图示约束。为避免展示错误关系,本次仅保留已验证的文字证据;重新评测会再次尝试生成。

补充参与方或组件说明谁参与、各自负责什么
说明关系与顺序提供输入输出、调用或依赖证据
重新评测自动选图按证据选择流程、时序或架构图
五维表现
该技能将主题转化为Vox风格视频,流程清晰,有示例和模型表,但依赖外部API和本地工具,安装与配置步骤基本完整,但缺少输出细节和错误处理说明。
质量证据
  • README中'How it works'部分描述了从主题到final.mp4的流程
  • 安装部分提供了git clone命令和API key设置
  • 模型表格列出了具体模型和用途
  • Quick start部分给出了示例提示词
  • Requirements部分列出了ffmpeg、Python等依赖
采用建议
优势
  • 问题与用途描述
  • 有效 README
  • 安装或接入步骤
  • 可执行示例
  • 未发现已知高风险模式
关注点
  • 缺少输出或结果说明
  • 缺少限制、权限或边界
  • 缺少错误处理或排障
  • 缺少输出格式和结果说明
  • 缺少错误处理和排障指南
适合

需要快速生成Vox风格解释视频的内容创作者、使用Claude Code或Codex等编码代理的开发者、有Atlas Cloud API密钥和ffmpeg环境的用户、希望自动化视频制作流程的团队

不建议直接用于

没有Atlas Cloud API密钥或ffmpeg的用户、需要完全本地化或离线运行的场景、对视频内容有严格版权或隐私要求的项目

也有自己的公开项目?先看完证据,再用当前规则生成独立报告。

评测我的项目 →
文档证据
71/100
问题与用途描述10 分
有效 README12 分
安装或接入步骤14 分
可执行示例16 分
输入、参数或工具说明11 分
输出或结果说明9 分
限制、权限或边界12 分
错误处理或排障8 分
许可证信息5 分
结构化章节3 分
安全证据
低风险
未发现已知高风险模式

静态扫描不是安全保证,生产接入前仍应人工复核权限和数据边界。

优先改进清单
  1. 01补充输出或结果说明
  2. 02补充限制、权限或边界
  3. 03补充错误处理或排障
方法、证据与局限展开
数据来源

GitHub Repository API

扫描范围

1 个文件 · 7,883 字符

评测引擎

v3.10.0 · AI 复核已启用(deepseek-chat)

局限
  • 静态评测不会安装或执行项目代码
  • 安全扫描基于高信号文件与已知模式,不能替代人工审计
  • 流行度只反映采用程度,不代表安全或工程质量

30 天热度趋势

README

English · 简体中文

🎬 Vox Director

Turn one topic into a finished Vox-style paper-collage explainer / ad video — script, collage keyframes, motion, voice-over, music and captions, all automated.

An agent skill that runs end to end on the Atlas Cloud API + local ffmpeg, usable by any coding agent (Claude Code, Codex, etc.). You give it a one-line topic; it gives you an mp4.

License: MIT Powered by Atlas Cloud Agent Skill

https://github.com/user-attachments/assets/ed08d230-7bcb-4b48-a17d-23c079208f9f

▶ "The evolution of Chinese civilization" · 30s

How football conquered the worldMexican street foodA brief history of moneyA brief history of Silicon Valley
Football history · 60sMexican street food · 60sA brief history of money · 60sSilicon Valley history · 60s

▶ more films — click any thumbnail to play


What it is

The look is the modern editorial paper-collage popularized by Vox explainers: hand-cut paper cut-outs, torn edges, tape, halftone dots, newspaper clippings, bold flat color per beat, big cut-out headlines — brought to life with motion, a narrator, music and captions.

How it works

One topic flows through one script per stage, all driven by a single beats.json per project:

topic
  │
  ├─ 1. beat map        pick a narrative arc → write beats.json      ◀── GATE 1: you approve the beat map
  ├─ 2. style bake-off  render the same beat in 3–4 themes           ◀── GATE 2: you pick the look by eye
  ├─ 3. keyframes       one collage poster per beat  (nano-banana-2)
  ├─ 4. motion          animate each poster          (gemini-omni-flash i2v)
  ├─ 5. voice + music   one narrator (xai/tts) + BGM (minimax/music)
  ├─ 6. assemble        ffmpeg: concat, duck music under VO, burn captions + watermark
  └─ final.mp4

That flow is B-roll — a topic in, everything generated. Two more input modalities reuse the same engine:

  • A-roll — you already have a talking-head video. It is ASR-segmented into beats and re-styled into the collage look, keeping the real face, lip-sync and gestures frame-for-frame (gemini-omni-flash/video-edit, auto-retrying on seedance-2.0/reference-to-video).
  • C-roll — you have one still photo (a selfie, a product shot). The subject is cut out as a photographic sticker — never redrawn — and each beat's poster is generated around it (nano-banana-2/edit). The narration can be cloned into the subject's own voice.

Two ideas make or break the result, and the skill is built around both:

  1. The look is born in the image step. Each beat is a finished collage poster. All the collage DNA (torn paper, cut-outs, halftone, headline text) lives in that image — if the poster isn't a rich collage, nothing downstream saves it.
  2. The motion is added after. By default an AI video model animates the whole poster (the "living poster" path). For dramatic piece-by-piece assembly, an optional local keyframe engine cuts the poster into parts and drives them frame-by-frame (no content filters, pixel-exact — great for real people).

Two human decision gates keep you in control (approve the beat map; pick the style); everything else is automated.

Models (verified on Atlas Cloud)

JobModel
Keyframe / collage postergoogle/nano-banana-2/text-to-image
Animate (non-real content)google/gemini-omni-flash/image-to-video
Animate (real people / brands)kwaivgi/kling-video-o3-pro/image-to-video
Re-style a talking-head (A-roll)google/gemini-omni-flash/video-edit
Anchor a photo in the collage (C-roll)google/nano-banana-2/edit
Narrationxai/tts-v1
Narration in a real person's voicebytedance/seed-audio-1.0 (voice cloning)
Musicminimax/music-2.6
Cut out an element (advanced path)youchuan/v8.1/remove-background

Model IDs drift — the skill fetches the live list from GET https://api.atlascloud.ai/api/v1/models before running.

Install

This is an agent skill — it works with any coding agent that can read a workflow and run scripts (Claude Code, Codex, …). Claude Code auto-discovers it as a skill; other agents read AGENTS.md → SKILL.md.

Option A — from this repo:

git clone https://github.com/Alisa0808/vox-director.git ~/.claude/skills/vox-director

Option B — from the packaged skill: download vox-director.skill and install it via your Claude skills UI.

Then set your Atlas Cloud API key (get one at atlascloud.ai/console/api-keys):

export ATLASCLOUD_API_KEY="sk-..."

Quick start

Just ask your coding agent, with the skill installed:

"Make me a Vox-style collage video introducing Mexican street food — English, 16:9, 15 seconds."

The agent will draft a beat map for your approval, run a style bake-off for you to pick from, then generate keyframes → motion → voice → music and assemble out/<project>/final.mp4.

Requirements

  • A coding agent — Claude Code, Codex, or similar
  • Atlas Cloud API key
  • ffmpeg + ffprobe (brew install ffmpeg)
  • Python 3 with Pillow (pip install pillow) — for caption/watermark overlays

What's in the box

SKILL.md              the skill (English) — the workflow the agent follows
SKILL.zh.md           the same skill in Chinese
AGENTS.md             entry point for non-Claude agents (Codex, …)
references/           the creative engine
  prompt-guide.md       the LOOK layer — prompt structures, vocab & 9 theme presets
  beat-layer.md         14 narrative arcs + hook/pacing + shot patterns
  voices.md             xai/tts voice roster — pick a voice_id per language/tone
  models-and-gotchas.md every API / ffmpeg gotcha, already solved
  local-engine.md       the advanced element-level motion engine
scripts/              one script per pipeline stage
examples/             ready-to-run beats.json examples
assets/               the showcase film

Credits

Built by @alisaqqt — follow for more agent-skill experiments.

Inspired by the collage-ad workflows of Stav Zilber, rom1trs and Higgsfield, and by Vox's explainer visual language.

Built end to end on Atlas Cloud — one prompt, one film.

License

MIT © 2026 Alisa Qian