Megadose AI progress, ranked and analyzed.

LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows

· HF Daily Papers ·
LynnReal-Omni is pitched as a single video model for controlled generation, editing, restoration, and long-form output from mixed visual inputs.

The paper says the framework uses a 32B shared multimodal diffusion transformer to combine text, images, references, 3D renders, and game recordings. It also introduces a 27B Flash version aimed at real-time rendering. The authors report 22-frame 540p generation and decoding on one H100 at 843 ms for LynnReal-Omni and 377 ms for Flash. They also describe a data pipeline and MSAVP, a 100-prompt, 20-metric evaluation setup for instruction following, plausibility, quality, temporal behavior, and audio coordination. HF Daily Papers' note

score 5

Categories: Research