2026 Beijing Academy of Artificial Intelligence Conference | Beyond the Debate Over Technical Routes, Spirit AI Bets on Real-World Model Evolution

2026 Beijing Academy of Artificial Intelligence Conference | Beyond the Debate Over Technical Routes, Spirit AI Bets on Real-World Model Evolution

From June 12 to 13, the 8th Beijing Academy of Artificial Intelligence (BAAI) Conference was held in Beijing, bringing together leading minds to explore frontier topics including embodied intelligence, world models, and self-evolving AI systems. Spirit AI made its appearance with the Moz1 embodied robot, showcasing generalized desktop organization and capsule toy machine tasks powered by its latest embodied model, Spirit v1.6, as well as a near-zero-latency teleoperation experience.
 

At the highly anticipated CEO roundtable featuring leading embodied AI companies, Spirit AI Founder and CEO Han Fengtao shared his strategic view that the industry should focus at its current stage on large-scale foundation model pre-training and the accumulation of high-quality data. Meanwhile, Co-founder and Chief Scientist Gao Yang joined the “Embodied Intelligence and Humanoid Robots” forum and delivered a keynote titled Reshaping the Physical World: Building a General-Purpose Robot Brain. He offered an in-depth look at the technical logic behind omni-modal input-output models and explored how closing the system loop could significantly reduce deployment costs across real-world scenarios.
 

As one of the few embodied AI teams in China with full-stack in-house capabilities spanning both AI algorithms and robotic hardware, Spirit AI has remained focused on three priorities: strengthening foundation models, building a scalable data system, and reducing deployment costs.
 

At BAAI Conference 2026, Spirit AI brought this vision to the center of the industry conversation.


Part 1

 

Moz Takes the Stage: Autonomous Reasoning, Precise Teleoperation

 

The Moz1 robot showcased at the conference is powered by Spirit AI’s proprietary Spirit v1.6 embodied foundation model.
 

By integrating world modeling with end-to-end control, Spirit v1.6 brings together multimodal perception, future world-state prediction, and fine-grained action generation, enabling robots to understand dynamic environments, interact autonomously with the physical world, and replan in real time.
 

Built on this core AI foundation, Moz1 moves beyond conventional preset rules and hard-coded logic, demonstrating a high degree of flexibility and adaptability. In the uncontrolled environment of the exhibition floor, Moz1 completed three live demonstrations.

 

Desktop Organization: Autonomous Long-Horizon Task Planning

 

Faced with a cluttered desk scattered with stationery, fruit, and drinks, Moz1 received just one simple instruction:

“Organize the desk.”
 

Powered by the end-to-end reasoning capabilities of Spirit v1.6, Moz1 autonomously completed the entire process, from high-level semantic understanding to complex task decomposition. Its “thought process” was displayed on the exhibition screen in real time: opening the drawer, placing objects, organizing pens, arranging fruit—with each step of the action sequence clearly visible.
 

Moz1 also demonstrated strong robustness and real-time replanning when faced with deliberately introduced disturbances.
 

After Moz1 placed the fruit in the designated area, a staff member moved it back out. Moz1 immediately detected the change, replanned its next move, and returned the fruit to the plate. When someone closed the drawer before Moz1 could execute the planned closing action, the robot briefly reassessed the environment, skipped the now-unnecessary step, and moved directly to the next task.
 

Rather than rigidly following a predefined action list, Moz1 continuously adapts its plan based on changes in the physical environment.
 

Beyond rigorous technical tests, Moz1—or “Xiao Mo”—also showed a more expressive side. When a staff member placed a flower on the desk, Moz1 smoothly recognized it, picked it up, and presented it to a guest. Through delicate force control at its fingertips, the robot turned a technical demonstration into a warm moment of human-robot interaction.

 

Capsule Toy Interaction: Robust Target Tracking Amid Disturbances







One of the most popular attractions at the conference, the capsule toy experience drew long lines soon after the exhibition opened.
 

Powered by end-to-end reasoning, Moz1 seamlessly completed the full sequence of turning the knob, retrieving the capsule, picking it up, and handing it to the participant, delivering surprise-filled capsule toys directly into visitors’ hands.
 

Technically, the number of turns required to dispense a capsule varies randomly each time. Instead of mechanically executing a preset motion sequence, Moz1 continuously monitors the dispensing slot through its vision system. The moment it detects a capsule dropping into the tray, it makes a closed-loop decision and immediately stops turning the knob. The entire process—from visual input to action output—is completed autonomously by the model.
 

The robotic arm then tracks and predicts the three-dimensional trajectory of the participant’s hands at high frequency. Even when visitors move their hands unpredictably, Moz1 can dynamically correct its motion path within milliseconds and smoothly deliver the capsule.

 

Teleoperation Experience: Near-Zero-Latency Motion Synchronization







The high-precision teleoperation zone was another highlight of the Spirit AI booth, shifting the experience seamlessly from autonomous robotic reasoning to intuitive human-robot synchronization.
 

Visitors could directly teleoperate Moz1 themselves. Powered by Spirit AI’s proprietary WBC (Whole-Body Control) algorithm, the system keeps human-robot motion latency at the millisecond level. Moz1 reproduces the operator’s hand movements in real time while automatically filtering out natural hand tremors during transmission.
 

Such fine-grained physical interaction places extremely high demands on control robustness in uncertain environments.
 

In the candied fruit skewer task, for example, the robot must steadily guide a bamboo skewer through fruit with relatively high resistance. Excessive force can damage the fruit. With WBC, Moz1 can automatically adjust the force applied based on fingertip feedback, enabling compliant and precise control.
 

The demonstration showcased the strength of Spirit AI’s underlying technology in handling high-frequency interactions and rapidly changing physical environments.


Part 2


Spirit AI CEO Han Fengtao: Strengthening Large-Scale Model Pre-training and Exploring High-Value Applications




On the morning of June 13, BAAI Conference hosted a dedicated CEO roundtable featuring leading embodied AI companies.
 

Spirit AI Founder and CEO Han Fengtao argued that the current gap between embodied model capabilities and deployment costs means the conditions for large-scale commercialization are not yet fully mature. The industry’s key priority this year should therefore be large-scale pre-training and the accumulation of high-quality data, laying the groundwork for the next true scaling phase.


Since its founding, Spirit AI has regarded high-quality data infrastructure as one of its core strategic priorities. The company has established more than 300,000 data collection points across China, supported by over 1,000 full-time data collection specialists.
 

Han Fengtao:

“The key constraint preventing embodied intelligent robots from achieving large-scale deployment is deployment cost. Today’s models may have the intelligence level of a one- or two-year-old child, which is why pre-training remains the primary focus of model development at this stage.”

“If a perfect humanoid robot represents a score of 100, today’s AI may only be at around 3. But with large models, the leap from 3 to 50 can happen extremely quickly. That is precisely why this industry is advancing at such remarkable speed.”

“Compute is no longer particularly scarce, and new model architectures and world-model technologies continue to emerge. But without sufficient data, even the best architecture cannot realize its full potential. From our perspective, data is the real bottleneck facing the industry today.”


Part 3


Spirit AI Co-founder & Chief Scientist Gao Yang: Closing the Software-Hardware Loop to Drive Down Model Deployment Costs




On the afternoon of June 13, Spirit AI Co-founder and Chief Scientist Gao Yang delivered a keynote titled Reshaping the Physical World: Building a General-Purpose Robot Brain at the “Embodied Intelligence and Humanoid Robots” forum, where he also participated in a panel discussion.
 

During his keynote, Gao introduced The Omni Model, Spirit AI’s proprietary omni-modal input-output model.
 

Rather than treating VLA and world models as competing technical routes, the framework unifies them as different objective functions within the same model. The model can simultaneously generate action sequences and predict future world states.
 

Gao also shared a quantitative finding from Spirit AI’s internal research: for every 10× increase in the amount of pre-training data used by the Spirit model family, the amount of post-training fine-tuning data required for downstream tasks decreases by 2.5×.
 

As pre-training continues to scale, the deployment cost of a new task could potentially fall from the current “10-hour scale” to just minutes.
 

During the panel discussion, Gao further outlined his view of the development path for embodied intelligence.
 

The industry, he argued, will progress through two stages. The first is a “bootstrap phase,” in which wearable data-collection devices and internet video are used to build large-scale pre-training corpora. The goal is to push foundation models toward a critical point where the marginal cost of deployment becomes extremely low, allowing virtually any new task to be deployed with only minimal fine-tuning.
 

Once this threshold is crossed, the industry will enter a second stage characterized by large-scale deployment, continuous real-world data feedback, and a self-reinforcing flywheel of model evolution.
 

Gao Yang:

“The debate over VLA versus world models is ultimately not that important. They are simply different objective functions serving the same goal of robotic manipulation. If both are useful, why not train them together within the same model?”

“Once we collect real-world data at sufficient scale, the marginal deployment cost of any new task will become extremely low. A small amount of fine-tuning will be enough.”

“Over the next 12 to 24 months, embodied foundation models will make tremendous progress. We are about to enter the beginning of a truly flourishing era.”

 

Part 4

 

Conclusion

 

Over the two-day conference, questions around deployment timelines, the relative importance of data and models, and the physical grounding of world models were repeatedly examined and debated.
 

Behind these discussions over technical paradigms lies a shared bet on the same future for embodied intelligence: bridging the gap between the virtual and physical worlds so that robots can truly evolve through interaction with reality.
 

This emerging industry consensus is also the path Spirit AI has been pursuing:
 

Build stronger foundation models. Build a scalable data system. Drive down deployment costs.
 

As a leading embodied AI team with full-stack in-house capabilities spanning embodied foundation models and robotic hardware, Spirit AI is leveraging its nationwide data collection network and the newly upgraded Spirit embodied foundation model to continuously narrow the gap between AI and physical reality.
 

We remain committed to advancing the generalization capabilities of embodied intelligence and enabling robotic technologies to achieve meaningful, large-scale deployment across a broader range of high-value industrial applications.