WorldExam

Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Yuxue Yang1,*,† Shuyao Shang1,* Jiahe Wang1 Zitong Zhou1 Liang Tan1
Junhan Zeng1 Ruizhi Li1 Junyan Li1 Yu Liu4 Xiao Yang5 Yong Li5
Jun Zhu5 Hongsheng Li2,3 Tieniu Tan1 Lue Fan1,†, Zhaoxiang Zhang1,

1 CASIA CASIA 2 SLAI SLAI 3 CUHK CUHK 4 AMAP AMAP 5 THU THU

* Equal Contribution Project Leaders Corresponding Authors

3
Control Interfaces
Camera-driven Action-driven Language-driven
4
Diagnostic Levels
Visual Quality Control Adherence Spatial Consistency World Reactivity
8
Evaluation Tasks
Camera Control Subject Control Scene Revisit Terrain Interaction Object Interaction Social Interaction Physical Reaction Goal Completion
1,474
Test Cases
First- and third-person viewpoints Human, animal, vehicle, and robot subjects Outdoor and indoor real scenes 3D renderings, cinematic footage, close-up views, animation, and dashcam videos

Overview

WorldExam is a hierarchical diagnostic benchmark for controllable video world models. Beyond visual appearance and explicit instruction following, it asks whether a model preserves a coherent world and exhibits inherent reactivity. Across camera-, action-, and language-driven interfaces, WorldExam organizes 1,474 test cases into four diagnostic levels and eight dedicated tasks under a unified evaluation protocol.

Inherent Reactivity: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input

WorldExam’s central diagnostic capability
WorldExam Overview

Visual Quality

Imaging Quality Photometric Consistency Temporal Flickering Subject Consistency Motion Smoothness Aesthetic Quality

Visual Quality measures the video's apparent appearance, including perceptual plausibility, temporal stability, and aesthetic quality.

Control Adherence

Camera Control Subject Control

Control Adherence measures whether the controlled camera or subject follows the input control.

Spatial Consistency

Scene Revisit 3D Consistency

Spatial Consistency measures whether the model preserves a coherent world when the camera revisits a previously observed viewpoint.

World Reactivity

Terrain Interaction Object Interaction Social Interaction Physical Reaction Goal Completion

World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input.

Comparison

Comparison with representative world-model benchmarks. The table compares supported model paradigms, viewpoints, task coverage, case counts, and evaluated models.

Benchmark Model
Paradigm
Viewpoint WorldExam Evaluation Tasks #Cases #Models
First
Person
Third
Person
Camera
Control
Subject
Control
Scene
Revisit
Terrain
Inter.
Object
Inter.
Social
Inter.
Physical
React.
Goal
Compl.
WorldScore C / L 3,000 20
MIND A 250 2
Omni-WorldBench C / L 1,068 18
WorldMark C / A / L 500 6
iWorld-Bench C / A / L 4,900 14
WBench C / A / L 289 20
WorldOlympiad A / L 1,000 8
WorldRoamBench A 600 10
WorldExam (Ours) C / A / L 1,474 20
Legend
C = Camera-driven models
A = Action-driven models
L = Language-driven models
= Supported / Evaluated
= Not supported / Not evaluated
FPV = First-person viewpoint
TPV = Third-person viewpoint
indicates that the instruction specifies the expected interaction consequence; the corresponding WorldExam tasks leave the evaluated reaction unstated.

Leaderboard

Select an evaluation track and model paradigm. Click column headers to sort the results. All retained metrics are normalized so that higher values indicate better performance.

Track:
Model Paradigm:

Static-Scene Track

The task scores diagnose Control Adherence and Spatial Consistency, whereas the general metrics characterize Visual Quality.

Task-Specific Metrics General Quality Metrics Aggregate Scores
# Model Average Task-Specific Metrics General Metrics
Overall Task General Camera
Control
Scene
Revisit
3D
Consistency
Photometric
Consistency
Temporal
Flickering
Aesthetic
Quality
Imaging
Quality
1 NeoVerse 85.39 93.29 77.49 97.33 89.25 98.19 70.83 93.53 52.72 72.19
2 WorldPlay 81.61 82.63 80.58 92.74 72.51 98.69 81.06 95.90 53.92 73.31
3 InSpatio-World (1.3B) 81.40 85.92 76.88 85.94 85.90 98.44 64.88 93.68 53.78 73.60
4 TrajectoryCrafter 78.00 81.68 74.31 80.32 83.03 96.98 62.83 92.56 52.17 67.03
5 Infinite-World 74.33 64.52 84.13 71.70 57.34 99.91 92.09 95.99 56.27 76.41
6 LingBot-World 70.39 58.27 82.50 58.19 58.35 99.59 81.52 96.53 59.87 75.00
7 Matrix-Game 3.0 69.28 70.30 68.26 76.36 64.25 95.59 30.14 93.40 48.80 73.39
8 Hailuo 2.3 68.64 55.99 81.28 63.29 48.70 99.38 79.88 95.33 56.30 75.52
9 ReCamMaster 67.15 53.33 80.97 38.64 68.01 99.33 81.97 95.52 54.58 73.45
10 Yume 1.5 66.02 54.06 77.97 75.67 32.45 98.44 67.53 95.18 53.42 75.28
11 Voyager 65.84 66.23 65.46 56.19 76.27 89.54 27.43 93.04 53.59 63.69
12 Wan 2.6 I2V 65.71 52.52 78.90 57.72 47.32 99.49 69.42 94.13 53.93 77.53
13 HappyHorse 1.0 65.32 50.37 80.27 58.29 42.45 99.62 74.17 95.02 55.56 77.00
14 Kling 2.5 63.98 44.50 83.46 50.18 38.81 99.87 87.87 97.56 56.28 75.70
15 Veo 3.1 60.99 41.94 80.03 40.83 43.05 99.06 73.61 95.43 55.49 76.57
16 Vidu Q3 60.24 40.90 79.57 42.75 39.05 99.43 71.28 95.01 55.38 76.73
17 Seedance 1.5 59.91 44.21 75.61 49.18 39.23 97.31 56.60 94.81 54.45 74.88
18 FantasyWorld 58.68 37.12 80.23 18.46 55.79 98.57 75.60 95.91 57.11 73.98
19 Hunyuan-GameCraft 57.59 41.55 73.62 41.33 41.77 93.61 53.93 93.91 54.97 71.69
20 Astra 57.16 35.37 78.95 32.59 38.15 96.46 78.41 96.38 51.71 71.78

Dynamic-Interaction Track

The dynamic-interaction track contains Subject Control and the five World Reactivity tasks and applies only to compatible action- and language-driven models; Goal Completion is language-only.

Task-Specific Metrics General Quality Metrics Aggregate Scores
# Model Average Task-Specific Metrics General Metrics
Overall Task General Subject
Control
Terrain
Interaction
Object
Interaction
Social
Interaction
Physical
Reaction
Goal
Completion
Subject
Consistency
Motion
Smoothness
Aesthetic
Quality
Imaging
Quality
1 Veo 3.1 72.77 65.02 80.52 37.28 44.71 75.96 85.10 61.76 85.30 92.07 99.12 58.22 72.65
2 Vidu Q3 72.35 64.18 80.51 27.67 64.39 71.59 81.91 61.23 78.26 92.78 98.76 57.24 73.25
3 Hailuo 2.3 72.03 63.37 80.69 36.49 61.57 67.01 72.45 63.84 78.86 93.25 99.31 57.74 72.47
4 HappyHorse 1.0 70.57 60.60 80.54 33.11 56.30 65.70 76.17 47.01 85.33 92.69 98.85 57.37 73.23
5 Wan 2.6 I2V 66.60 52.29 80.91 29.02 49.21 44.56 66.40 48.15 76.39 94.46 98.15 56.30 74.72
6 Seedance 1.5 66.50 53.36 79.64 32.51 53.83 37.91 72.09 47.60 76.24 92.14 98.87 56.72 70.84
7 LingBot-World 60.76 39.91 81.61 55.47 24.33 25.94 60.37 33.43 94.98 98.89 60.86 71.69
8 Kling 2.5 60.45 39.85 81.04 28.40 35.95 27.70 66.80 31.99 48.25 96.00 99.48 56.86 71.83
9 WorldPlay 57.57 37.86 77.28 49.75 27.49 33.75 51.40 26.91 88.06 98.09 54.19 68.76

Samples

Representative cases across the eight evaluation tasks are shown below.

Select First Frame

Video Comparison

Citation

If you find our work useful, please consider citing:

@article{yang2026worldexam,
  title   = {WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity},
  author  = {Yang, Yuxue and Shang, Shuyao and Wang, Jiahe and Zhou, Zitong and Tan, Liang and Zeng, Junhan and Li, Ruizhi and Li, Junyan and Liu, Yu and Yang, Xiao and Li, Yong and Zhu, Jun and Li, Hongsheng and Tan, Tieniu and Fan, Lue and Zhang, Zhaoxiang},
  journal = {arXiv preprint arXiv:2608.02603},
  year    = {2026}
}