Hardcore Long-Horizon Tasks × Adorably Approachable! Spirit AI Steals the Spotlight at WAIC as Moz2 Wins Over the Crowd on Its Public Debut

Hardcore Long-Horizon Tasks × Adorably Approachable! Spirit AI Steals the Spotlight at WAIC as Moz2 Wins Over the Crowd on Its Public Debut

On July 17, the 2026 World Artificial Intelligence Conference (WAIC) officially opened in Shanghai. Spirit AI’s next-generation commercial service robot, Moz2, made its first public appearance. With its friendly, approachable design and forward-looking AI-native product philosophy, Moz2 quickly became one of the most popular attractions in the embodied intelligence exhibition area, marking a new generation of Spirit AI products designed for commercial service scenarios.
 

At this year’s exhibition, Spirit AI focused on two key themes—expanding the boundaries of real-world applications and validating technological capabilities through deployment—and presented two major developments:

Moz2, Spirit AI’s next-generation commercial service robot, made its public debut, marking the company’s official expansion into commercial service applications including retail, hospitality, and office buildings, and extending the deployment boundaries of general-purpose embodied intelligence.
 

Spirit v1.6 showcased an upgraded spatial-level, long-horizon compound task, enabling Moz1 to autonomously complete a coherent sequence of multiple steps from a single instruction—“Tidy up the living room.” The demonstration addresses a common industry limitation in which robots rely on short-term memory and fragmented task execution.
 

From real-world deployment in industrial environments to frontier exploration in commercial services, Spirit AI is advancing embodied intelligence across diverse physical environments along a clear roadmap: “Industry First, Commercial Services Next, Home Last.”


Moz2 Makes Its Commercial Service Robot Debut
Expanding into a New Frontier of Service Robotics


As Spirit AI’s first exploratory product for the commercial service sector, Moz2 became an undeniable crowd favorite at this year’s exhibition. Unlike the industry’s traditional “build the hardware first, add intelligence later” development approach, Moz2 was designed from the outset around an AI-native philosophy, serving as a key platform for extending Spirit AI’s general-purpose embodied intelligence into commercial service scenarios.
 

Approachable Design That Stands Out: Moz2 features a soft cream-beige body with rounded contours, paired with a custom scarf that gives it a warm and approachable visual identity. On-site, it supports natural voice conversations and physical interactions, with smooth, fluid movements. On the opening day alone, it attracted large crowds of visitors stopping for photos and interactions, quickly becoming one of the exhibition area’s most popular attractions.
 

AI-Native Co-Design: Moz2 adopts an AI-native co-design approach integrating the “embodied intelligence system + robot body + uDAS data collection device,” enabling architectural alignment between the robot hardware and the data collection system. Human operation data can be efficiently transferred directly to the robot body, while also laying the foundation for rapidly transferring Moz1’s mature long-horizon planning and fine manipulation capabilities to Moz2, enabling general-purpose capabilities to be reused across different products.


Modular Hardware Ecosystem: Moz2 features 32 degrees of freedom, an omnidirectional steering-wheel chassis, and a foldable leg structure, together with 7-DoF biomimetic arms and 4-DoF three-finger dexterous hands, balancing mobility with manipulation potential. Multiple magnetic expansion points are built into the body, allowing its appearance to be quickly customized for different brands and scenarios, while functional accessories such as delivery racks and inspection modules can also be attached. This highly reusable modular design enables Moz2 to adapt to diverse commercial service requirements.
 

Multimodal Safety System: Designed for public environments with dense pedestrian traffic, Moz2 will be equipped with a multimodal perception system combining LiDAR, vision, and ultrasonic sensing. Together with compliant control algorithms, the system enables high-precision mapping and dynamic obstacle avoidance, helping ensure reliable operation in environments where humans and robots coexist.


At this stage, Moz2 will focus primarily on refining human-robot interaction and validating user experience in commercial service scenarios. Going forward, supported by Spirit AI’s general-purpose embodied intelligence technology stack, Moz2 will progressively inherit the full-stack intelligent capabilities already validated at scale on Moz1, while continuing to evolve across applications such as retail shelf stocking, hotel delivery, building inspection, and exhibition hall guidance.
 

Spirit v1.6 Powers Long-Horizon Compound Tasks

Moz1 Takes on the “Tidy Up the Living Room” Challenge Live

 

At this year’s exhibition, Spirit AI’s self-developed Spirit v1.6 embodied foundation model presented four real-robot task demonstrations spanning four key dimensions: long-horizon compound tasks, dynamic planning, fine manipulation, and general-purpose interaction, comprehensively showcasing full-stack technical capabilities that have been validated in real-world industrial environments. Among them, the newly unveiled long-horizon living-room tidying task was a major highlight of Spirit AI’s booth and represented a significant technological breakthrough for the embodied AI industry.
 

Tackling the Industry Bottleneck of Long-Term Memory:

One Instruction to Tidy Up the Living Room

 

Today, the embodied AI industry continues to face a common challenge of “short-term memory and segmented execution.” Most embodied foundation models suffer from a “goldfish memory” problem: they can respond only to isolated, short-duration instructions, but struggle to carry out multi-step, long-horizon tasks continuously. During execution, they can easily forget previous objectives or lose task context, creating a major barrier to bringing general-purpose robots from laboratory demonstrations into complex real-world environments.
 

Building on its existing long-horizon task capabilities, Spirit AI has now extended long-horizon intelligence to spatial-level, unstructured environments. With just one natural-language instruction—“tidy up the living room”—Moz1 can autonomously complete four consecutive subtasks, including putting away Coke cans and placing dirty dishes into the dishwasher, without requiring humans to break the task down into individual steps.
 

The scenario recreates a real-world unstructured environment at a 1:1 scale, with object positions and other environmental conditions kept random and dynamic. During the demonstration, staff can also randomly toss additional objects, such as crumpled paper, onto the table. The robot can perceive these environmental changes in real time, automatically incorporate the newly introduced objects into its task sequence, and place them in the trash bin, demonstrating strong robustness against disturbances.


360 Group Founder Zhou Hongyi Interacts with Moz1 On-Site
 

Dynamic Task Replanning: Visualizing the Full Decision-Making Pipeline


In the tabletop organization task, when the environment is deliberately disrupted, the robot can automatically identify changes in the scene and reorganize its task sequence, demonstrating the embodied foundation model’s strong generalization and adaptability in dynamic, unstructured environments.
 

Fine Manipulation: Millimeter-Level Grasping for Precise Sorting





The pill-sorting demonstration focuses on fine manipulation of small objects. Using vision to accurately identify the position and shape of individual pills, the robot precisely controls grasping force and placement accuracy to rapidly sort and arrange multiple types of pills, demonstrating close coordination between the model algorithms and robot control system.


General-Purpose Interaction: Precise Manipulation of a Capsule Toy Machine


The capsule toy machine demonstration further validates the robot’s general-purpose manipulation capabilities. The robot can autonomously identify the position of the machine’s knob, precisely control the force and angle of rotation, and reliably complete the entire process of dispensing a capsule toy.




Underlying these capabilities is Spirit AI’s self-developed Spirit v1.6 embodied foundation model. The model adopts a deeply integrated architecture combining VLA and world models. Unlike the modular pipelines commonly used across the industry, it connects environmental perception, task understanding, action planning, and state prediction into an end-to-end decision-making pipeline, enabling a human-like process of perceiving, deciding, and adapting continuously. This gives the system stronger generalization and robustness in complex, dynamic environments.




This general-purpose technology stack, already validated at scale in industrial scenarios, forms the foundation for Spirit AI’s continued expansion into new application domains. It also connects the underlying logic behind the coordinated evolution of Moz1 and Moz2. Moz1’s deep deployment in high-end manufacturing environments, including CATL, provides high-quality real-world training data for continuous model iteration, further improving the generalization and robustness of the Spirit embodied foundation model. Meanwhile, Moz2’s exploration of commercial service scenarios such as retail and hospitality is building practical experience for bringing general-purpose robots into a broader range of service environments. Sharing the same underlying technology stack, Moz1 and Moz2 create a two-way reinforcement between technological capabilities and application boundaries—an important step in Spirit AI’s long-term strategy of “industry first, commercial services next, and home applications later.”
 

Embodied AI is now at a critical stage of transitioning from technological validation to large-scale deployment. Looking ahead, Spirit AI will continue to place technological innovation at the center of its development. On the one hand, it will further deepen large-scale deployment in industrial scenarios and strengthen the technological foundation of its general-purpose embodied intelligence. On the other hand, it will steadily advance product refinement and real-world exploration in commercial service scenarios, driving embodied AI to evolve from a “tool that executes instructions” into an “intelligent service agent capable of adapting to diverse scenarios,” while contributing to the global competitiveness of China’s embodied AI industry and the intelligent transformation of the real economy.