NewsPronto

 
The Times


.

News from Asia

HiDream Unveils HiDream-O1-Video-1.0, a Native Omnimodal Video Model Built for Physical Consistency

  • Written by Media Outreach

The model supports multimodal inputs including text, images and video, and generates high-fidelity 1080p videos of 5 to 20 seconds with natively synchronized audio

BEIJING, CHINA – Media OutReach Newswire – 17 September 2026 – HiDream.ai, an AI company specializing in foundation models and generative AI, announced the launch of HiDream-O1-Video-1.0, or HiDream V1, its native omnimodal video generation model. image Designed around a deeper understanding of creative intent and real-world physics, HiDream V1 supports multimodal inputs including text, images and video. It can generate high-fidelity 1080p videos ranging from 5 to 20 seconds, with enhanced narrative planning, character consistency, physical plausibility and audiovisual synchronization. In its debut on two independent international benchmarks, HiDream V1 ranked No. 4 globally on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai Image-to-Video leaderboard, placing it among the world's leading video generation models.
A Strong Debut on Two International AI Benchmarks
Artificial Analysis independently evaluates leading AI models through standardized and reproducible benchmark testing. Arena.ai uses anonymous head-to-head comparisons and user voting to assess model outputs. Together, the two platforms provide third-party perspectives on model performance based on both standardized testing and real-world user preferences. HiDream V1's results demonstrate its competitiveness across key dimensions including visual quality, prompt adherence, motion quality, narrative coherence and audiovisual coordination. The global AI video generation market has become increasingly competitive, with models developed by Chinese teams—including Seedance and MiniMax H3—maintaining strong positions on international leaderboards. HiDream V1's top-tier debut further expands China's presence among the world's leading video generation models. "The next generation of video models will not be defined solely by higher resolution or longer duration," said Yao Ting, Chief Technology Officer of HiDream.ai. "What matters is whether a model can genuinely understand a creator's intent and how objects, actions and sounds interact in the real world. HiDream V1 was designed from the outset to represent text, video and audio within a unified framework. Our goal is to move video generation beyond simply looking sharp and moving smoothly toward understanding instructions, sustaining coherent performances and keeping sound aligned with visuals. Its performance on two independent international benchmarks provides encouraging validation of our native omnimodal approach." Upgrades in Visual Fidelity, Narrative Coherence and Character Consistency HiDream V1 introduces broad improvements in high-fidelity rendering, narrative continuity and character consistency. Rather than optimizing only for the quality of individual frames, the model is designed to maintain coherence across an entire video sequence. Once characters, emotions, motivations and physical rules are placed within a continuous sequence, each element can affect the others. HiDream therefore uses a technical framework built around three stages: planning first, followed by joint generation, and then alignment through multimodal reward signals. Under this approach, the model first plans the narrative and character states at a global level. It then jointly constrains visuals, movement and semantics during generation, before using multimodal reward signals to align visual quality, continuity and physical plausibility. Native Omnimodal Architecture for Better Intent and Physics Understanding The central challenge in AI video generation is shifting from the quality of individual frames to a broader understanding of complex instructions, continuous motion and the rules governing the physical world. User prompts often contain multiple layers of information, including characters, actions, settings, camera directions, dialogue and sound. To handle these requirements more effectively, HiDream V1 incorporates multimodal intent understanding and planning. Before generation begins, the model structures the user's request across elements such as shot duration, setting, character state, movement, facial expression, composition, camera motion, dialogue and ambient sound. Through this "understand, plan and generate" workflow, HiDream V1 plans narrative development and character states at a global level, while coordinating visuals, motion, semantics and audio during generation. This is designed to improve content completeness, character consistency and continuity between shots. Physical reasoning is another core capability of HiDream V1. The model incorporates factors such as gravity, inertia, collisions, deformation, materials, lighting and spatial continuity into the generation process. As a result, object movement, character actions and environmental responses can more closely...

Read more: HiDream Unveils HiDream-O1-Video-1.0, a Native Omnimodal Video Model Built for Physical Consistency