StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
Researchers have introduced StemVLA, an open-source vision-language-action model designed to improve robotic manipulation through 3D spatial awareness and 4D historical data integration. By incorporating future geometry knowledge and past temporal context, the model aims to enhance how robots interpret their surroundings and execute complex physical tasks.
Covered by 1 source
- AarXiv CS.AI↗Jiasong Xiao, Yutao She, Kai Li, Yuyang Sha, Ziang ChengJul 1