Beyond Scale: Embodied Generalization Is a State Representation Problem
July 9, 2026
Current embodied models, including recent VLAs (Vision-Language-Action models), still struggle when deployed in unseen environments or on unseen embodiments. They often look robust in benchmark settings, but fail once scene layout, object set, camera viewpoint, control frequency, or robot kinematics shift.
