Text-to-Motion Generation on KeyframeFace (test)
0.0348FIDExpress4D-MDM
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Express4D-MDMArchitecture=Transformer-based denoising diffusion, Dataset=KeyframeFace (adapted from Express4D)2025.12 | 0.0348 | 0.023 | 8.77 | 16.23 | 23.44 | 1.13 | 1.2555 | |
| Tuned-semantic-14BRepresentation=Semantic (P_sem), Training Status=Fine-tuned (LoRA), Scale=14B, Backbone=Qwen3-14B / DeepSeek-R1-Distill-Qwen-14B2025.12 | 0.0599 | 0.0145 | 20.79 | 34.13 | 45.79 | 0.87 | 1.0602 | |
| Tuned-semantic-4BRepresentation=Semantic (P_sem), Training Status=Fine-tuned (LoRA), Scale=4B, Backbone=Qwen3-4B-Instruct2025.12 | 0.0777 | 0.0177 | 19.95 | 34.13 | 43.99 | 0.83 | 1.0689 | |
| Tuned-non-semantic-4BRepresentation=Numerical (non-semantic), Training Status=Fine-tuned (LoRA), Scale=4B, Backbone=Qwen3-4B-Instruct2025.12 | 0.2406 | 0.031 | 9.74 | 18.87 | 25.48 | 0.63 | 1.1779 | |
| Tuned-non-semantic-14BRepresentation=Numerical (non-semantic), Training Status=Fine-tuned (LoRA), Scale=14B, Backbone=Qwen3-14B / DeepSeek-R1-Distill-Qwen-14B2025.12 | 0.2805 | 0.0323 | 10.46 | 19.23 | 25.72 | 0.67 | 1.1773 | |
| Base-semantic-4BRepresentation=Semantic (P_sem), Training Status=Pre-trained (Base), Scale=4B, Backbone=Qwen3-4B-Instruct2025.12 | 0.8557 | 0.0585 | 8.77 | 14.66 | 20.91 | 1.26 | 1.322 | |
| Base-non-semantic-4BRepresentation=Numerical (non-semantic), Training Status=Pre-trained (Base), Scale=4B, Backbone=Qwen3-4B-Instruct2025.12 | 4.3623 | 0.1796 | 3.25 | 6.61 | 9.5 | 2.95 | 1.4691 |