nvidia/ARDY-G1-RP-25FPS-Horizon8
Captured source
source ↗ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
[Paper](https://research.nvidia.com/labs/sil/projects/ardy/assets/ardy_paper.pdf), [Project Page](https://research.nvidia.com/labs/sil/projects/ardy/)
Description:
ARDY is an autoregressive diffusion model designed for interactive motion generation, supporting online text prompting and flexible long-horizon kinematic constraints (root paths/waypoints, full-body keyframes, and sparse joint positions/rotations) with real-time responsiveness.
ARDY-G1-RP-25FPS-Horizon8 was developed by NVIDIA as a part of the ARDY project. It was trained on the Bones Rigplay 1 dataset with the 34-joint Unitree G1 robot skeleton at 25 fps. See [below](#model-versions) for other model variants.
This model is ready for commercial or non-commercial use.
License/Terms of Use:
Use of this model is governed by the NVIDIA Open Model Agreement
Deployment Geography:
Global
Use Case:
Developers and researchers with any level of animation experience can use ARDY to generate controllable humanoid motions in their real-time applications. This could include motion planning for humanoid robots, character movement in digital twin and industrial simulations, digital human motion for synthetic data, and animations for games and other interactive applications.
Release Date:
HuggingFace: 07/10/2026 via HuggingFace
Reference:
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
Model Architecture:
Architecture Type: Diffusion Model
Network Architecture: Novel Two-Stage Transformer
Number of model parameters: 326 M
Input:
Input Type(s): Text, Other: Pose Constraints, History Poses
Input Format(s): String, Tensor
Input Parameters: One-Dimensional (1D), N-Dimensional (ND)
Other Properties Related to Input: History pose duration is max 8 sec.
Output:
Output Type(s): Other: Pose Sequence
Output Format: Tensor
Output Parameters: N-Dimensional (ND)
Other Properties Related to Output: Pose sequence contains global root translation and joint rotations. Output poses have max duration of 8 sec.
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
Software Integration:
Runtime Engines: PyTorch
Supported Hardware Microarchitecture Compatibility:
- NVIDIA Ampere
- NVIDIA Blackwell
- NVIDIA Hopper
Supported Operating System(s): Linux
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
Model Versions:
- ARDY-Core-RP-20FPS-Horizon40: 27-joint "core" skeleton at 20 fps, 40 frame generation horizon
- ARDY-Core-RP-20FPS-Horizon8: 27-joint "core" skeleton at 20 fps, 8 frame generation horizon
- ARDY-G1-RP-25FPS-Horizon52: 34-joint Unitree G1 robot skeleton at 25 fps, 52 frame generation horizon
- ARDY-G1-RP-25FPS-Horizon8: 34-joint Unitree G1 robot skeleton at 25 fps, 8 frame generation horizon
This repo corresponds to the ARDY-G1-RP-25FPS-Horizon8 model variant. Please refer to the codebase for installation and usage instructions.
Training, Testing, and Evaluation Datasets:
The model was trained and evaluated using the Bones Rigplay 1 dataset.
Training Dataset:
Data Modality:
- Text
- Other: Human Motion Capture
Text Training Data Size: Less than a Billion Tokens
Other Training Data Size: 630 hours of human motion captures
Data Collection Method by dataset: Automatic/Sensors
Labeling Method by dataset: Hybrid: Automated/Human
Properties: Contains optical motion capture data with corresponding text descriptions covering a diverse range of behaviors such as locomotion, everyday activities, and gestures. Motions are clipped to 10 sec long and resampled to the desired FPS for training. An LLM is used to augment the dataset with diverse paraphrases of text labels.
Testing Dataset:
Data Collection Method by dataset: Automatic/Sensors
Labeling Method by dataset: Hybrid: Automated/Human
Properties: 70 hours of motion data held out from training. The test split contains motions from content categories not seen in training.
Evaluation Dataset:
Benchmark Score: See codebase for evaluation results.
Data Collection Method by dataset: Automatic/Sensors
Labeling Method by dataset: Hybrid: Automated/Human
Properties: Same as test dataset.
Inference:
Acceleration Engine: TensorRT
Test Hardware:
- NVIDIA A100
- NVIDIA RTX 4090
Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
For more detailed information on ethical considerations for this model, please see the Model Card++ Bias, Explainability, Safety & Security, and Privacy Subcards below.
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
Bias
Field | Response :---------------------------------------------------------------------------------------------------|:--------------- Participation considerations...
Excerpt shown — open the source for the full document.