Vision-language-action models: foundation models that map camera images and language instructions directly to robot actions, trained on large teleoperation datasets.
What is it connected to
Related
Instancesπ0PRIMARY · High · 2 sourcesHelixPRIMARY · High · 1 source