machinelearning.apple.com·8d ago
REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs
Submitted by @

Most current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they generate raw motor commands or very short sequences of actions, without organizing behaviors into reusable, well-defined abstractions. As a result, these models perform poorly on long-horizon (multi-step) tasks, and it’s difficult to interpret what they have learned. Existing approaches for discovering skills often avoid the core problem of deciding when two action sequences a
Apple Machine Learning Research
Published 8d ago ago
Original reporting by Apple Machine Learning Research. GridIndex is an aggregation and intelligence layer — full credit to the original publisher.
This page summarizes and tracks coverage of this developing story.
Read full story at Apple Machine Learning Research 0 views 0 upvotes 0 comments 0 shares
Discussion · 0
Sign in to join the discussion.