Sequential Text-to-Motion · BABEL
Long-form human motion composition from an ordered sequence of text prompts. Every method is retargeted to the same SMPL-22 joints; uTMR FID is computed only after per-sample L2 normalization of the motion embeddings.
Normalized uTMR FID
Compare all 1,295 episodes
BABEL validation split
1,295 validation episodes
30 fps
Motius Joint-Position Evaluator
L2-normalized embedding FID
0
Ranked methods
1,295
Episodes
7,285
Captioned segments
5,990
Transition boundaries
Method Comparison
Generated methods only; reference rows never set the chart scale.
Normalized Profile
100 is the best generated-method value on each axis.
All-Case Comparison
Search every processed validation episode, switch between all measured methods, and inspect the exact caption active at each frame.
Results
Click any metric header to sort. Best and second-best values are recomputed within the active view.
BestSecondReference, excluded from ranking
Protocol
A single conversion and metric contract for every sequential-generation method.
- Prompts
- All 1,295 eligible processed BABEL validation episodes; 7,285 action intervals after short-action merging and LLM caption rewriting.
- References
- Paired GT subsequences and transition windows from the same episode boundaries; no independent reference pools.
- Representation
- Neutral zero-beta SMPL-22 joints66. Every semantic subclip is independently first-frame canonicalized before uTMR; each transition window is canonicalized once so its cross-boundary gap remains intact.
- Semantic Metrics
- Official BABEL act_cat action-group R-Precision, nearest-positive MM-Dist, and L2-normalized embedding FID from the Motius Joint-Position Evaluator. Synonymous captions are retained without becoming false negatives.
- Transition Metrics
- L2-normalized embedding FID and AUJ gap are ranked; Diversity and absolute Peak Jerk are diagnostic reference statistics.
- Ranking
- GT/reference and paper-only values are visible but cannot receive best or second-best styling.