NewsPronto

 
The Times


.

News from Asia

HiDream.ai Launches HiDream-O1-Embodied, Extending Its Native Omni-Modal World Model Strategy into Physical Interaction

  • Written by Media Outreach

HiDream-O1-Embodied tops RoboColiseum’s Robustness leaderboard, highlighting the model’s ability to maintain stable performance under complex real-world conditions

BEIJING, CHINA - Media OutReach Newswire - 8 September 2026 - HiDream.ai has officially launched HiDream-O1-Embodied, an embodied world model designed to advance physical interaction for embodied intelligence. Built on HiDream.ai's native omni-modal technology strategy, the model enhances robots' physical perception, dynamic prediction, and execution capabilities, enabling more robust interaction with the physical world. image The launch marks an important step in HiDream.ai's broader effort to connect image, video, 3D, and action modalities within a unified architecture. By extending its world model capabilities from understanding and reasoning to action and execution, HiDream.ai is building a closed-loop technical foundation for native omni-modal intelligence. Alongside its release, HiDream-O1-Embodied made its debut on RoboColiseum, an embodied intelligence model evaluation platform. The model ranked No. 1 on the platform's Robustness leaderboard, achieving an average score of 0.692. "We believe a complete world model foundation requires three core capabilities: omni-modal representation, causal reasoning, and physical-world modeling — all centered on the ability to express, understand, and generate within the real world," said Ting Yao, CTO of HiDream.ai. "From the beginning, HiDream.ai's native omni-modal world model architecture was designed to support unified representations across modalities, including action. The release of HiDream-O1-Embodied marks a critical milestone in our technology roadmap, as we move from simulating the world to enabling AI to operate in the real world." HiDream-O1-Embodied Tops RoboColiseum's Robustness Leaderboard with a Score of 0.692 RoboColiseum is a standardized simulation benchmark for embodied intelligence models, designed to provide a multidimensional and reproducible evaluation framework. Through high-fidelity simulation tasks that closely approximate real-robot performance, the platform helps developers assess model strengths and limitations while continuously tracking progress across the field. Open to universities, research institutions, model developers, and researchers worldwide, RoboColiseum continuously updates its evaluation results with the goal of establishing a reliable benchmark for embodied models. Built on high-fidelity simulation environments that closely mirror real-world conditions, RoboColiseum evaluates models across four major dimensions: instruction following, spatial understanding, robustness, and general-purpose manipulation. These dimensions are assessed through four capability leaderboards and 78 high-fidelity simulation tasks. Since entering internal testing, RoboColiseum has attracted dozens of leading models from China and abroad. Among its evaluation dimensions, Robustness is widely regarded as one of the most challenging. It measures a model's stability and generalization under non-ideal conditions by varying backgrounds, lighting, materials, robot initial states, camera positions, and image quality, while also introducing diverse paraphrases of instructions. In other words, this is not a test conducted in the "greenhouse" of a lab environment. It is designed to evaluate how well a model performs when faced with the kinds of uncertainty, variation, and interference that robots are likely to encounter in the real world. HiDream-O1-Embodied ranked first on the Robustness leaderboard with a score of 0.692, supported by HiDream.ai's native omni-modal foundation. A native omni-modal world model provides an inherent basis for cross-modal understanding, generation, and action. At the execution level, HiDream-O1-Embodied introduces advances across three core capabilities, enabling more precise instruction understanding, more reliable perception, and stronger resistance to environmental interference. Language Understanding: Moving Beyond Keyword Matching Traditional robots often interpret language instructions at the level of keyword matching. A robot may understand "Bring me the cup," but change the phrasing to "Get me a cup" or "Hand me the cup," and it may fail to respond correctly. HiDream-O1-Embodied covers an equivalent instruction space encompassing diverse verbs, sentence structures, and expressions. Rather than being constrained by specific wording, it focuses on the underlying intent. No matter how an instruction is phrased, the model can move beyond the literal wording and accurately identify what the user actually means. Visual Perception: Multi-View Collaboration for Greater Reliability In the physical world, a robot's visual input is rarely ideal. Camera positions may shift, calibration accuracy can change over time, and individual visual feeds may be obstructed or disrupted. HiDream-O1-Embodied integrates information from multiple viewpoints,...

Read more: HiDream.ai Launches HiDream-O1-Embodied, Extending Its Native Omni-Modal World Model Strategy into...