FLUX 3 Action: A 7B World Action Model for Robot Control
The open-weight 7B World Action Model jointly predicts future video and robot actions, achieving up to a 42.2% success rate on RoboLab.
- Single-step distillation achieves a 38.3% success rate on RoboLab while running 1.34x to 2.28x faster per second of motion than Pi0.5 in FP8 on workstation and datacenter GPUs.
- Guidance-distilled checkpoints achieve 42.2% on RoboLab, outperforming Cosmos 3 Nano (36.8%) while delivering 1.52x to 3.95x speedups in FP8.
- The model predicts a 2.13-second motion horizon compared to 1.0s for Vision-Language-Action baselines like Pi0.5.
- Hybrid delegation with GPT 6 Astra achieves a 90% success rate on complex tasks while cutting costs to $8.77 and runtimes to 8 minutes per success.
- Pretrained on video, image, and audio data, with midtraining across gaming inputs, egocentric hand poses, and teleoperation datasets.
Robotics developers get an open-weight world action model that achieves state-of-the-art policy accuracy while remaining fast enough for local deployment.

Sources
Read this as text
Back to the AI news