Sequential Text-to-Motion · BABEL

Long-form human motion composition from an ordered sequence of text prompts. Every method is retargeted to the same SMPL-22 joints; uTMR FID is computed only after per-sample L2 normalization of the motion embeddings.

Normalized uTMR FID Compare all 1,295 episodes
BABEL validation split 1,295 validation episodes 30 fps Motius Joint-Position Evaluator L2-normalized embedding FID
0
Ranked methods
1,295
Episodes
7,285
Captioned segments
5,990
Transition boundaries

Method Comparison

Generated methods only; reference rows never set the chart scale.

Ranked Values

Normalized Profile

100 is the best generated-method value on each axis.

All-Case Comparison

Search every processed validation episode, switch between all measured methods, and inspect the exact caption active at each frame.

Results

Click any metric header to sort. Best and second-best values are recomputed within the active view.

BestSecondReference, excluded from ranking

Protocol

A single conversion and metric contract for every sequential-generation method.

Prompts
All 1,295 eligible processed BABEL validation episodes; 7,285 action intervals after short-action merging and LLM caption rewriting.
References
Paired GT subsequences and transition windows from the same episode boundaries; no independent reference pools.
Representation
Neutral zero-beta SMPL-22 joints66. Every semantic subclip is independently first-frame canonicalized before uTMR; each transition window is canonicalized once so its cross-boundary gap remains intact.
Semantic Metrics
Official BABEL act_cat action-group R-Precision, nearest-positive MM-Dist, and L2-normalized embedding FID from the Motius Joint-Position Evaluator. Synonymous captions are retained without becoming false negatives.
Transition Metrics
L2-normalized embedding FID and AUJ gap are ranked; Diversity and absolute Peak Jerk are diagnostic reference statistics.
Ranking
GT/reference and paper-only values are visible but cannot receive best or second-best styling.