Meta tags:
description= Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control;
keywords= ;
Headings (most frequently used words):
steerable, policies, to, embodied, reasoning, hierarchical, control, commands, vision, language, action, for, and, can, flexibly, follow, diverse, allowing, them, better, interface, with, vlms, transfer, foundation, model, capabilities, the, real, world, abstract, experiments, bibtex, steering, generating, data, robot, in, context, learning,
Text of the page (most frequently used words):
the (33), and (30), #steerable (22), #policies (22), commands (19), steering (16), level (15), language (11), for (11), embodied (11), context (11), learning (11), robot (11), reasoning (10), vlms (9), vlm (9), vision (8), hierarchical (8), control (8), vlas (8), action (7), this (7), our (7), two (7), approach (6), are (6), task (6), methods (6), high (6), with (6), models (6), diverse (6), trained (5), over (5), command (5), styles (5), into (5), low (5), how (4), allow (4), which (4), model (4), standard (4), that (4), tasks (4), train (4), using (4), capabilities (4), foundation (4), based (4), can (4), appropriate (3), abstraction (3), past (3), where (3), vla (3), fine (3), physical (3), semantic (3), prior (3), novel (3), controlling (3), openvla (3), policy (3), reasoner (3), carrot (3), better (3), real (3), world (3), manipulation (3), synthetic (3), use (3), data (3), move (3), grasp (3), these (3), thus (3), from (2), william (2), chen (2), jagdeep (2), bhatia (2), catherine (2), glossop (2), nikhil (2), mathihalli (2), ria (2), doshi (2), andy (2), tang (2), danny (2), driess (2), karl (2), pertsch (2), sergey (2), levine (2), used (2), select (2), only (2), modalities (2), find (2), outperforms (2), many (2), like (2), examples (2), behaviors (2), also (2), leverage (2), behavior (2), its (2), need (2), perform (2), multi (2), towel (2), make (2), grounded (2), put (2), both (2), dataset (2), instantiate (2), experiments (2), automatically (2), pipeline (2), various (2), them (2), left (2), gripper (2), open (2), container (2), motions (2), subtasks (2), instructions (2), levels (2), natural (2), however (2), often (2), full (2), off (2), shelf (2), abstractions (2), allowing (2), steerability (2), introduce (2), pretrained (2), across (2), knowledge (2), reason (2), interface (2), generalization (2), website, borrowed, under, creative, commons, attribution, sharealike, international, nerfies, article, chen26, title, author, year, 2026, bibtex, uniquely, enabled, one, baseline, restricted, subtask, confirming, performance, benefits, prompting, saycan, paraphrased, correct, erroneous, stalling, apply, grained, infer, most, reasons, style, observes, resulting, iteratively, refines, adaptively, improve, casts, obviating, structured, scene, representations, rely, step, stack, pots, blue, block, object, plate, tune, autoregressively, generates, rationale, explaining, what, should, before, picking, execute, when, either, equivalent, non, ablation, effective, generalizable, watermelon, pot, adapting, codebases, study, ways, application, bridge, widowx, openpi, acquire, attaching, labels, existing, trajectories, stage, extract, relevant, features, then, query, api, compile, all, gemini, parse, demonstrations, segments, labeled, generating, hybrids, along, traces, above, pointing, atomic, reach, set, spanning, beyond, include, descriptions, performed, too, induce, range, skills, necessary, solving, formulaic, vague, illustrative, adapt, experience, tuning, produces, chain, thought, rationales, decompose, showcase, new, highlighting, enables, hierarchial, architectures, training, label, demos, automated, annotation, core, limitation, transfering, robotics, limited, detailed, visual, inferences, settings, providing, valuable, common, sense, priors, robotic, effectively, grounding, remains, challenge, employ, executed, separate, between, usually, fundamentally, limits, much, steer, rich, pixel, coordinates, improving, controllability, unlock, enabling, improved, demonstrate, benefit, learned, prompted, via, extensive, outperform, baselines, including, challenging, long, horizon, abstract, flexibly, follow, transfer, narrated, preview, muted, code, paper, intelligence, stanford, university, berkeley,
Text of the page (random words):
steerable vision language action policies for embodied reasoning and hierarchical control steerable vision language action policies for embodied reasoning and hierarchical control william chen 1 jagdeep bhatia 1 catherine glossop 1 nikhil mathihalli 1 ria doshi 2 andy tang 2 danny driess 3 karl pertsch 3 sergey levine 1 1 uc berkeley 2 stanford university 3 physical intelligence paper code model data preview muted full narrated steerable policies can flexibly follow diverse commands allowing them to better interface with vlms to transfer foundation model capabilities to the real world abstract pretrained vision language models vlms can make semantic and visual inferences across diverse settings providing valuable common sense priors for robotic control however effectively grounding this knowledge in robot behaviors remains an open challenge prior methods often employ a hierarchical approach where vlms reason over high level commands to be executed by separate low level policies e g vision language action models vlas the interface between vlms and vlas is usually natural language task instructions which fundamentally limits how much vlm reasoning can steer low level behavior we thus introduce steerable policies vlas trained on rich synthetic commands at various levels of abstraction like subtasks motions and grounded pixel coordinates by improving low level controllability steerable policies can unlock pretrained knowledge in vlms enabling improved task generalization we demonstrate this benefit by controlling our steerable policies with both a learned high level embodied reasoner and an off the shelf vlm prompted to reason over command abstractions via in context learning across extensive real world manipulation experiments these two novel methods outperform prior embodied reasoning vlas and vlm based hierarchical baselines including on challenging generalization and long horizon tasks steerable policies our steerable policies are vision language action models trained on diverse and detailed steering commands a core limitation of transfering foundation model capabilities to robotics is limited low level policy steerability we thus introduce steerable policies vision language action models vlas trained on diverse steering commands for training data we use an automated annotation pipeline to label robot demos with synthetic steering commands we instantiate two steerable policies based on the openvla and π 0 5 architectures we showcase two new hierarchial control methods highlighting how steerability enables better use of vlm capabilities fine tuning a vlm into an embodied reasoner which produces chain of thought rationales for how to decompose tasks into appropriate steering commands using an off the shelf vlm to perform robot in context learning over steering abstractions allowing it to adapt its commands based on past experience steering commands illustrative examples of steering command styles used to train steerable policies standard vlas are trained on task level commands i e high level natural language descriptions of the task to be performed however these commands are often too formulaic and vague to induce the full range of physical skills necessary for solving novel manipulation tasks we thus train our policy on steering commands a diverse set of instructions spanning many styles and levels of abstraction beyond standard task level commands we also include semantic subtasks e g reach for the carrot and grasp the container atomic motions e g move left and grasp pointing e g open gripper above the container at x y gripper traces e g move along x 1 y 1 x 2 y 2 hybrids of these styles e g move left from x 1 y 1 to x 2 y 2 to grasp the carrot generating data we use foundation models to automatically parse robot demonstrations into segments that are labeled with diverse steering commands to train steerable policies we need a dataset of steering commands we acquire this by automatically attaching synthetic language labels to existing robot trajectories using a multi stage pipeline we leverage various foundation models to extract relevant embodied features then query an api based vlm gemini to compile them into commands of all our steering styles hierarchical control experiments our two hierarchical control methods using steerable policies we instantiate two steerable policies by adapting both the openvla and π 0 5 openpi codebases and train it on the bridge widowx dataset using this vla we study two ways in which steerable policies allow better application of vlm capabilities to real world manipulation embodied reasoning put the carrot in the pot put the watermelon on the towel controlling steerable policies with high level embodied reasoning vlms is an effective approach for generalizable control we fine tune a vlm into a high level embodied reasoner that autoregressively generates a grounded rationale explaining what the robot should do before picking a steering command to execute with the low level vla when controlling either the openvla or π 0 5 steerable policy we find that this approach outperforms equivalent standard vlas past embodied reasoning methods and a hierarchical non reasoning ablation robot in context learning make the blue block the only object on the plate stack the pots on the towel steerable policies allow high level vlms to perform robot in context learning on novel multi step tasks steerable policies also allow vlms to leverage in context learning where the model reasons to select a steering command style observes the resulting behavior and iteratively refines its commands to adaptively improve on the task this approach casts robot in context learning as standard vision language in context learning obviating the need for structured scene and action representations that prior robot in context learning methods rely on paraphrased examples of how in context learning over steering commands allow vlms to correct erroneous or stalling behaviors apply fine grained physical and semantic reasoning and infer which command styles are most appropriate as in context learning is used to select the appropriate level of abstraction for steering the robot this approach is uniquely enabled by our steerable policies as past vlas are only trained on one or two steering modalities we find our approach outperforms a saycan like baseline where the vla is restricted to subtask level commands confirming the performance benefits of in context learning over many prompting modalities bibtex article chen26 steerable policies title steerable vision language action policies for embodied reasoning and hierarchical control author william chen and jagdeep bhatia and catherine glossop and nikhil mathihalli and ria doshi and andy tang and danny driess and karl pertsch and sergey levine year 2026 website borrowed from nerfies under a creative commons attribution sharealike 4 0 international
|