📍 ECCV 2026 Workshop

Human Motion-Informed World Models &
Socially Intelligent Action

Bridging human motion modeling, visual world models, and embodied AI for socially intelligent perception and action.

📅 September 9, 2026 · 08:00 CEST · Morning Session (AM)
📍 Room Malmömässan D2
🏃 Motion Modeling 🌍 World Models 🤖 Embodied AI 🚗 Autonomous Systems
Workshop Schedule

Half-day program · Morning of September 9, 2026

📅 Wednesday, September 9, 2026 · 08:00–12:30 📍 Room Malmömässan D2
08:00–08:30

Welcome & Opening Remarks

08:30–09:10

Keynote 1 — Matthieu Cord

End-to-end Models for Autonomous Driving · View speaker details

09:10–09:50

Keynote 2 — Gerard Pons-Moll

Towards Compositional, Embodiment-Agnostic, Streaming and Interactive Human Motion Generation · View speaker details

09:50–10:30

Coffee Break & Poster Session

10:30–11:10

Keynote 3 — Jakob Engel

From Egocentric Perception to Human-Centric World Models · View speaker details

11:10–11:50

Keynote 4 — Misa Komuro

From Predicting Humans to Cooperating with Humans · View speaker details

11:50–12:20

Panel Discussion

Open Q&A with speakers: Gerard Pons-Moll, Jakob Engel, and Misa Komuro

12:20–12:30

Workshop Best Paper Award & Closing Remarks

Keynote Speakers

Speaker biographies are available below. Talk titles and abstracts will be added as they are confirmed.

Matthieu Cord

Matthieu Cord

Full Professor, Sorbonne University
Director, valeo.ai
World models, vision-language models
Keynote Talk

End-to-end Models for Autonomous Driving

Abstract

Autonomous driving is shifting from modular pipelines toward holistic end-to-end architectures powered by foundation models. This talk presents DrivoR, a camera-only end-to-end planner that compresses sensor data into a few "scene tokens" and separates trajectory generation from trajectory scoring, enabling controllable driving behavior with a compact and scalable design. We then look beyond imitation learning: test-time trajectory optimization to surpass expert demonstrations, and ongoing work on world-model-based planning, where a learned dynamics model and reward model allow the agent to imagine and evaluate the consequences of its actions before acting.

Biography

Matthieu Cord is a Professor at Sorbonne University and Director of valeo.ai. He is also Principal Investigator of the Large Generative Vision-Language Models Chair at the Sorbonne Center for Artificial Intelligence. His research focuses on large-scale representation learning, vision-centric world models, and vision-language models.

Gerard Pons-Moll

Gerard Pons-Moll

Full Professor
University of Tübingen
Human motion, 3D tracking, virtual humans
Keynote Talk

Towards Compositional, Embodiment-Agnostic, Streaming and Interactive Human Motion Generation

Abstract

We want to generate human motion that does exactly what we ask: for animation, for populating simulation, for training data, and eventually for robots. Today we ask with language, and language is a coarse interface. A sentence cannot say which part of the body moves, whose body it is, what happens next, or what the body is touching. I will argue that progress now depends on adding those four things back, and describe recent work on each: part-level compositional control, motion representations shared across arbitrary skeletons and morphologies, streaming generation that plans ahead and runs at interactive rates, and generation grounded in contact with the scene. I will close on where this leads, since the same fine-grained control over virtual humans is what is needed to control robots acting in the real world.

Biography

Gerard Pons-Moll is a Professor in the Department of Computer Science at the University of Tübingen. He is core faculty at the Tübingen AI Center and leads the Real Virtual Humans research group. His research covers human motion understanding and modeling, 3D tracking and reconstruction, analysis of people in videos, and virtual human models.

Jakob Engel

Jakob Engel

Sr. Director of Research
Meta Reality Labs
SLAM, egocentric perception, spatial AI
Keynote Talk

From Egocentric Perception to Human-Centric World Models

Abstract

Egocentric devices like Project Aria are incredibly capable perception devices; capable of tracking motion at millimeter accuract, recover how a person moves — pose, hands, gaze — and reconstruct the space around them. This talk connects those perception capabilities to learning human-centric world models that represent the distribution of how humans move and interact with the world and each other. Such models, which span both human and robot data and are built from and for embodied, multi-modal data from purpose-built hardware rather than data scraped from the internet, represent the next major shift for computer vision and robotics.

Biography

Jakob Engel is Sr. Director of Research at Meta Reality Labs, where he leads egocentric machine perception and Spatial AI research for Project Aria. He received his PhD in Computer Science from the Technical University of Munich in 2016. His work on direct visual odometry, including DSO and LSD-SLAM, received the ECCV 2024 Koenderink (test of time) award, and his research focuses on grounding AI models in the physical 3D world and enabling them to learn from humans changing that world.

Misa Komuro

Misa Komuro

Chief Engineer
Honda R&D
Autonomous driving, human-robot collaboration
WaPOCHI micro-mobility robot
WaPOCHI micro-mobility robot
Keynote Talk

From Predicting Humans to Cooperating with Humans

Abstract

As autonomous systems move from structured environments into everyday human spaces, their challenge is no longer simply to perceive the world and predict what will happen next. They must operate in environments where people continuously adapt to one another, where intentions are often implicit, and where the behavior of the autonomous system itself influences how others respond. This talk explores the challenges of building autonomous systems that can interact naturally and effectively with people, drawing on experiences from the development of WaPOCHI, a walking-support robot designed to operate in close proximity to humans. Real-world deployments have revealed that robust perception and motion prediction alone are often insufficient for seamless interaction. Instead, successful collaboration requires understanding behavior in the broader context of goals, intentions, and ongoing social dynamics. Recent advances in world models provide a promising foundation for addressing these challenges by enabling autonomous systems to represent and anticipate the evolution of complex environments. Building on these developments, the talk presents a broader vision of Cooperative Intelligence: autonomous agents that not only predict the world, but continuously understand, adapt to, and coordinate with the people and agents around them. Ultimately, the talk argues that the next frontier of autonomy lies not in prediction alone, but in enabling agents to participate as cooperative partners within human-centered environments.

Biography

Misa Komuro is Chief Engineer in the Innovative Research Excellence at Honda R&D Co., Ltd. She leads the development of WaPOCHI, a walking-support robot designed to accompany users, carry their belongings, and assist them in navigating crowded pedestrian environments.

About the Workshop

Human motion and activity provide a critical signal for understanding, predicting, and interacting with dynamic environments. In recent years, computer vision has made significant progress in human motion perception, activity understanding, and motion generation, providing a strong foundation for modeling human behavior and dynamics. Integrating these advances into visual world models that support prediction, planning, and decision-making is a key next step toward enabling embodied intelligent systems to reason and act effectively in human-populated scenes.

This workshop focuses on the challenge of integrating rich models of human motion and behavior into world models, an increasingly important yet still under-explored direction in visual world modeling. Topics include: 1. modeling human motion and activity in complex, interactive scenes; 2. learning dynamic world models that incorporate human behavior to represent scenes, objects, and affordances; 3. enabling efficient and robust real-world deployment, including integrated perception-planning, safe navigation and autonomous driving, and improved generalization under noise, occlusions, and distribution shifts.

This topic is closely aligned with recent progress in visual world modeling and generative simulation which aim to capture vision-based representations of the scene structure, object relations, and human dynamics. The workshop will bring together research on human motion and activity modeling, dynamic scene understanding, and world models that explicitly account for human behavior as a central component of the environment. By uniting perspectives from computer vision, embodied AI, robotics, and graphics, the workshop provides a forum to explore human-centered world models that enable socially intelligent perception and action in applications such as dynamic scene understanding, predictive navigation, autonomous driving, and human-robot interaction.

Topics of Interest
🏃‍♂️

Human Motion & Activity Modeling

  • Trajectory, pose, mesh, and flow representations of human motion
  • Vision-based motion perception, tracking, and forecasting
  • Human motion generation and digital human modeling
🌐

Dynamic World Models

  • Generative and predictive world models
  • Visual world models incorporating human motion and behavior
  • Scene, object, and affordance modeling conditioned on human motion
  • Scene representations for dynamic environments
🚀

Efficient and Robust Real-World Deployment

  • Integrated perception and planning in human motion-informed models
  • Design choices under computational and real-time constraints
  • Robustness to noise, occlusions, and distribution shifts
  • Generalization across environments, agents, and human behaviors
Submission Guidelines

Submission Tracks

We welcome both archival and non-archival submissions. All submissions will use the official ECCV format, up to 7 pages excluding references.

Archival

Original, unpublished work. Accepted archival papers will be included in the official ECCV 2026 Workshop Proceedings.

Non-Archival

Work-in-progress, preliminary results, or relevant work previously published. These will be presented at the workshop but will not appear in the proceedings.

Formatting and Review

The review process is double-blind: author identities will not be visible to reviewers, and reviewer identities will not be visible to authors. Please ensure your manuscript is properly anonymized.

How to Submit

All papers should be submitted through OpenReview. Please select "Archival" or "Non-Archival" as the submission type when submitting your paper.

Important Dates

Paper Submission Deadline
Jul 26, 2026 (Anywhere on Earth)
Notification of Acceptance
Aug 8, 2026Aug 10, 2026
Camera-Ready Deadline
Aug 15, 2026 (Anywhere on Earth)
Workshop Day
Sep 9, 2026 (Morning Session)
Accepted Papers

15 papers accepted · poster stands 186–200

Coffee Break & Poster Session · 09:50–10:30 · September 9, 2026
Poster #186

HOI-JEPA: Auditing Action Contribution in Interaction-Centric World Models

Yunze Liu; Li Yi

View on OpenReview →
Poster #187

Towards Human Motion World Models via Executable Behaviour Representations

Rimvydas Rubavicius; manisha dubey; Siddharth N; Subramanian Ramamoorthy

View on OpenReview →
Poster #188

Parkour-Sim: Controlled Simulation of Extreme Human Motion

Dena Bazazian; Marius N. Varga; Francisco Ramirez-Javega

View on OpenReview →
Poster #189

Context-Aware Adaptation of Pre-trained GNNs for Skeleton-Based Soccer Action Understanding

Yizhou Xu; Lars Bretzner; Tiesheng Wang; Atsuto Maki

View on OpenReview →
Poster #190

SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation

Shuang Liang; Jing He; Lejun Liao; Chuanmeizhi Wang; Guo Zhang; Ying-Cong Chen; Yuan Yuan

View on OpenReview →
Poster #191

Map-Mono-Ego: Map-Grounded Global Human Pose Estimation from Monocular Egocentric Video

Hiroyuki Deguchi; Ryosuke Hori; Tsubasa Maruyama; Kotaro Amaya; Mitsunori Tada; Hideo Saito

View on OpenReview →
Poster #192

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

Enrico Pallotta; Sina Mokhtarzadeh Azar; Lars Doorenbos; Serdar Ozsoy; Umar Iqbal; Juergen Gall

View on OpenReview →
Poster #193

Structured Natural Language as a Representation for Human Hand Gestures

Elizabete Munzlinger; Lucas Frey Torres Hanson; Mads Aqqalu Roager; Ted Vucurevich; Dan Witzner Hansen; Fabricio Batista Narcizo

View on OpenReview →
Poster #194

MuCHeR: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking

Guénolé Fiche; Philippe Weinzaepfel; Romain Brégier; Fabien Baradel

View on OpenReview →
Poster #195

EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion

Cecilia Curreli; Florian Hofherr; Dominik Muhle; Abhishek Saroha; Riccardo Marin; Daniel Cremers

View on OpenReview →
Poster #196

CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments

Tan-Dzung Do; Tuan Dat Phuong; Cuc T.Trinh; Vien Anh Ngo; An Thai Le

View on OpenReview →
Poster #197

Real Motion in Virtual Crowds: Quantifying Error Propagation in Human Motion Understanding Pipelines

David Anderlohr; Mickael Cormier; Jürgen Beyerer

View on OpenReview →
Poster #198

Benchmarking Motion Tokenizers: A Controlled Study of Human Motion Quantization

Kaiki Camilo de Oliveira; Iago Alves Brito; Lucas Arruda Fernandes; Walcy Rios; Arlindo Rodrigues Galvão Filho

View on OpenReview →
Poster #199

PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments

Chaeyun Kim; Kihyun Kim; Jongho Shin; Taeyoun Kwon; Junghyun Kim; Mijin Koo; Haon Park

View on OpenReview →
Poster #200

JVC: Joint Vision Cross-Attention for Fine-Grained Badminton Stroke Recognition

Navneeth Dhamotharan; Bin Han

View on OpenReview →
Organizers
PhD Student Organizers
Yasamin Borhani

Yasamin Borhani

EPFL, Switzerland

3D localization & trajectory forecasting

Mariam Hassan

Mariam Hassan

EPFL, Switzerland

World models for autonomous driving

Yang Gao

Yang Gao

EPFL, Switzerland

Motion prediction & 3D perception

Yi Yang

Yi Yang

KTH / Scania, Sweden

Prediction & planning in traffic

Yufei Zhu

Yufei Zhu

Örebro University, Sweden

Spatio-temporal data modeling

Senior Organizers
Andrey Rudenko

Andrey Rudenko

Postdoc, TU Munich, Germany

Motion prediction & HRI

Wanyu Ma

Wanyu Ma

Postdoc, The Chinese University of Hong Kong, China

Robotic manipulation

Achim J. Lilienthal

Achim J. Lilienthal

Full Professor & MIRMI Deputy Director, TU Munich, Germany

Multi-modal perception & navigation

Martin Magnusson

Martin Magnusson

Full Professor in computer science, Örebro University, Sweden

3D mapping & autonomous systems

Chuan Guo

Chuan Guo

Senior Research Scientist, Meta Reality Labs, USA

Generative AI for digital humans

Narūnas Vaškevicius

Narūnas Vaškevicius

Senior research scientist, Bosch Center for Artificial Intelligence, Germany

Dynamic perception & 3D scene graphs

Luigi Palmieri

Luigi Palmieri

Group leader and research scientist, Bosch Center for Artificial Intelligence, Germany

Predictive navigation & RL

Important Dates

All deadlines are Anywhere on Earth (AoE) time

Paper Submission
Jul 26, 2026
Notification
Aug 8, 2026Aug 10, 2026
Camera-Ready
Aug 15, 2026
Workshop Day
Sep 9, 2026
Morning (AM) session