Automatically translated.View original post

Practical etiquette to compare and verify multiple AI models in 30 minutes

"When I made it with Model A, the movement was unnatural, but when I remade it with Model B, the delivery date was exceeded by one day. I don't know what the correct answer is."

When I received this consultation from a student at the video school, I first checked. "Did you do a comparison test before the performance?" The answer was No.

Runway Gen-4.5 has a feature called "Aleph." It is a mechanism that can execute multiple models such as Kling, WAN 2.2, Seedance 2.0 in parallel on the same workspace and compare and verify with the same prompt and the same reference image.

The procedure is simple. Enable Multi-Model Mode and select up to three comparison models (about 5 minutes). Enter the same prompt and reference image for all models (about 10 minutes). Evaluate the naturalness of body movement, stability of background, and reproducibility of movement instructions to decide on the material to use (about 15 minutes). Model selection is completed within 30 minutes in total.

Recording "which model was good" on a project-by-project basis will continue to improve the accuracy of tool selection in similar projects. Changing trial and error into records will become a knowledge asset in the field.

▼ Full implementation procedure of Aleph function is here

https://note.com/videolife/n/na41a2a8f18cb?sub_rt=share_sb

# video _ production AI Videos The Runway # Video Production Workflow # Technical Director# GenerativeAI # VideoProduction

5/28 Edited to

... Read moreAI動画制作の現場で、似たような複数のAIモデルをどう評価すればよいか迷うことはよくあります。私自身も過去に、どのモデルが本当に最適か判断できず時間を浪費した経験があります。そんなときにRunwayのAleph機能に出会い、その威力を実感しました。 このAleph機能は、同じワークスペース上で複数のAIモデルを一度に走らせ、同条件で直接比較できるため、効率的に検証が進みます。私の場合、最初にMulti-Model Modeを有効化し、Kling、WAN2.2、Seedance 2.0などのモデルをセットアップ。約5分で設定を済ませた後、一つのプロンプトと参照画像を各モデルに入力しました。 その後、比較対象の動画を3つ並べて身体の動きの自然さ、背景のブレや安定性、動き指示がどれだけ忠実に再現されているかを目視で評価。評価基準をチームで共有しながら進めることで、主観的な判断を減らし、選定の質を高められました。 特にNEON TOKYOのような都市のネオンを用いた映像では、背景の安定性が仕上がりに大きく響くため、この比較方法で時間短縮とクオリティ確保の両立が可能になりました。また、こうした比較検証の結果をプロジェクト単位で記録しておくことで、次の案件で似た条件のモデル選定がスムーズにできるようになりました。 私の実体験として、この30分程度の比較検証を怠ると、後からの手戻りや納期遅延が発生しやすいです。逆にAleph機能を活用して定量的かつ効率的に評価を行えば、安心して制作を進められ、チームの信頼も得られます。 複数モデルの検証はAI動画制作のクオリティを左右する重要なプロセス。この記事で紹介された手法を取り入れることで、実務効果を大幅に高められると思います。より詳しいAleph機能の設定方法は記事のリンクも参考に、ぜひ試してみてください。