How to work through it

01

Freeze the system under test

Record provider, interface, region, account, model label, checkpoint or API path, repository revision, Context-IR use, regeneration use, date, duration, ratio, resolution, and audio setting. A local H3-Base run is not directly equivalent to a complete hosted 2K path. If the system changes during the evaluation, version the result instead of pooling outputs from different configurations into one score.

02

Use briefs that expose production risk

Prepare a small authorized set: a text-only action with camera direction; a first-to-last-frame transition; a subject or product reference under motion; a reference-led performance with dialogue or sound; and a delivery-specific portrait or landscape composition. Each brief needs acceptance criteria for meaning, identity, geometry, motion, timing, text, sound, continuity, safety, rights, and editability. Do not select only showcase-friendly prompts.

03

Measure decisions and revision effort

For each attempt, log whether the result passed, why it failed, what changed next, and how much human work remained. Useful measures include accepted outputs per fixed attempt budget, repeated failure categories, time spent preparing references, number of isolated revisions, audio repair, cleanup, edit integration, and review escalations. Publish the rubric and sample size beside any aggregate rather than presenting a context-free percentage.

04

Make a bounded selection decision

Separate documented facts, direct observations, reviewer judgments, and unresolved questions. A strong result may still fail because the workflow lacks the required control, predictable access, rights clearance, audit trail, or delivery setting. Recommend H3 only for the tested use, configuration, and review process. State the trigger for retesting, such as a model update, new checkpoint, pricing change, interface change, license revision, or new delivery requirement.