Meta tags:
description= Instruction following policies can be autonomously improved by a foundation model powered data collection system and learning algorithm.;
keywords= Autonomous Improvement, Instruction Following Skills, Scaled Data Collection;
Headings (most frequently used words):
soar, autonomous, improvement, data, of, instruction, following, skills, via, foundation, models, overview, reformulating, language, conditioned, control, results, quality, bibtex,
Text of the page (most frequently used words):
the (56), and (30), data (30), language (21), policy (21), soar (19), #autonomous (18), conditioned (17), #improvement (15), with (14), collected (12), collection (10), from (9), can (9), learning (8), goal (8), this (7), susie (7), instruction (6), skills (6), subgoal (6), trajectories (6), scenes (6), for (6), image (6), following (5), models (5), more (5), images (5), instructions (5), across (5), lcbc (5), gcbc (5), autonomously (5), then (5), that (5), self (5), diffusion (5), scale (5), dataset (4), than (4), vlm (4), generated (4), task (4), success (4), objects (4), used (4), training (4), internet (4), into (4), goals (4), are (4), robotic (3), human (3), both (3), diverse (3), low (3), different (3), not (3), pre (3), these (3), decomposed (3), level (3), its (3), test (3), better (3), ability (3), hindsight (3), formulation (3), robot (3), which (3), model (3), end (3), control (3), vlms (3), guide (3), interesting (3), improve (3), under (2), via (2), foundation (2), zhiyuan (2), zhou (2), pranav (2), atreya (2), abraham (2), lee (2), homer (2), walke (2), oier (2), mees (2), sergey (2), levine (2), arxiv (2), also (2), contribution (2), consists (2), each (2), trajectory (2), during (2), compared (2), datasets (2), much (2), time (2), weeks (2), minimal (2), effort (2), loop (2), will (2), useful (2), berkeley (2), 000 (2), over (2), top (2), quality (2), side (2), out (2), ground (2), interactions (2), contrast (2), high (2), generalizing (2), right (2), average (2), direct (2), trained (2), same (2), results (2), suboptimal (2), system (2), collect (2), semantically (2), relevant (2), enables (2), multiple (2), such (2), leverage (2), cheap (2), supervised (2), after (2), process (2), tasks (2), components (2), towards (2), semantic (2), help (2), understanding (2), supervision (2), use (2), large (2), without (2), website, template, borrowed, creative, commons, attribution, sharealike, international, license, nerfies, article, zhou2024autonomous, title, author, journal, preprint, 407, 20635, year, 2024, bibtex, release, secondary, our, deployment, comes, annotations, commanded, one, episode, label, predicted, similar, size, other, current, but, smaller, frames, matter, failure, contains, includes, publicly, available, cost, arm, hope, resource, offline, reinforcement, research, see, how, download, https, github, com, rail, transitions, sets, table, setups, bounds, magnitude, achievable, above, video, depicts, run, against, slightly, distribution, environment, some, present, like, purple, eggplant, lemon, were, seen, bridge, unable, containing, leading, them, exhibits, significantly, meaningful, attributed, degree, generalizability, novel, robust, grounded, generation, generalizes, well, thanks, left, single, generalist, all, boosts, rate, investigate, whether, truly, necessary, explore, compare, behavior, cloning, while, improves, performance, achieves, attribute, turn, optimal, objectives, reaching, observe, leads, evaluate, capabilities, evaluating, learn, reflects, effectiveness, procedures, being, tested, has, been, occur, find, context, very, semantics, separated, motor, allowing, former, latter, utilizes, denser, signal, would, require, separate, relabel, truth, purely, objective, decompose, style, editing, commands, converted, rolled, fixed, number, timesteps, new, repeats, until, instructpix2pix, reformulating, lapse, videos, depicted, running, succeeds, where, once, failed, putting, together, becomes, deployed, fleet, widowx, robots, enabling, within, just, few, evaluated, achieve, build, off, knowledge, stored, automate, synthesize, yet, detection, proposers, decouples, benefit, any, entirely, algorithm, namely, relabeled, generator, decouple, benefits, propose, particular, call, multi, key, ideas, standard, paradigm, improving, policies, involves, manually, collecting, additional, labelling, finetuning, expensive, hard, huge, instead, deploy, fleets, overview, code, paper, equal,
Text of the page (random words):
autonomous improvement soar autonomous improvement of instruction following skills via foundation models zhiyuan zhou pranav atreya abraham lee homer walke oier mees sergey levine uc berkeley equal contribution paper code dataset overview the standard paradigm for improving instruction following policies involves manually collecting additional robot data labelling it with language instructions and then finetuning the policy on this data this process is expensive and hard to scale up without huge human effort can we instead deploy robot fleets to collect large scale datasets without human supervision and use that data to self improve the policy we propose a particular formulation of an autonomous improvement loop which we call soar that enables self improvement of a multi task language conditioned policy the key ideas are use vlms and diffusion models to help guide large scale autonomous data collection decouple language understanding from robotic control so semantic understanding benefits from internet scale pre training and low level control improve with self supervision soar decouples a language conditioned policy into an image goal conditioned policy and a language conditioned image subgoal generator the benefit of such a formulation is that any autonomously collected data can be used to improve the policy with an entirely self supervised learning algorithm namely hindsight relabeled goal conditioned learning autonomous data collection can then build off of the internet scale knowledge stored in vlms and diffusion models vlms can be used as task proposers to guide the policy towards semantically interesting goals and can automate the success detection of autonomously collected trajectories diffusion models can be used to synthesize interesting and diverse image goals from the semantic goals generated by the vlm both of these components help guide data collection towards interesting yet diverse tasks putting these components together soar becomes an end to end system for autonomous improvement we deployed soar on a fleet of 5 widowx robots enabling the collection of over 30 000 autonomous trajectories across more than 50 scenes within just a few weeks and evaluated its ability to achieve a 2x improvement in multiple language skills across 10 test scenes after running soar the policy succeeds at tasks where it once failed time lapse videos of autonomous data collection generated subgoal images during data collection are depicted in the top right reformulating language conditioned control we decompose soar s instruction following policy into an image goal conditioned policy and susie an instructpix2pix style language conditioned image editing diffusion model language task commands from the vlm are converted into subgoal images with the diffusion model after which the goal conditioned policy is rolled out for a fixed number of timesteps then a new subgoal is generated with the same language instruction and the process repeats until the end of the trajectory in the context of autonomous improvement such a formulation is very useful semantics are separated from motor skills allowing the former to leverage cheap internet data and the latter cheap autonomously collected robot data goal conditioned learning utilizes a denser learning signal than language conditioned learning and can better leverage suboptimal autonomous data and the goal conditioned policy can be trained with a purely self supervised objective in contrast to a direct language conditioned policy which would require a separate model to hindsight relabel autonomous trajectories with the ground truth language improvement results to evaluate the improvement capabilities of soar we test the system on 10 different scenes evaluating its ability to autonomously collect semantically relevant data and then learn from it improvement reflects the effectiveness of both the learning and data collection procedures if data relevant to the skills being tested has not been collected improvement will not occur we find that soar enables on average a a 2x improvement in multiple language skills across each of the 10 scenes we observe that more autonomous data leads to better improvement training a single generalist policy on all the autonomously collected data across the 10 test scenes boosts the average success rate from 58 to 65 we then investigate whether the decomposed language conditioned policy gcbc susie is truly necessary to explore this we compare it with a direct language conditioned behavior cloning lcbc policy trained on the same autonomous data collected by gcbc susie while lcbc also improves performance the decomposed policy in soar achieves much better results 65 vs 28 we attribute this to goal conditioned learning s ability to turn suboptimal autonomous trajectories into optimal objectives for reaching hindsight goals autonomous data quality left lcbc data collection right soar s gcbc susie data collection the quality of the collected autonomous data bounds the magnitude of improvement achievable the above side by side video depicts a data collection run with lcbc as the instruction following policy compared against the gcbc susie policy used in soar on this slightly out of distribution environment some of the objects present like the purple eggplant and lemon were not seen in the pre training dataset bridge v2 and so lcbc is unable to ground language instructions containing these objects leading to minimal interactions with them in contrast the data collected by gcbc susie exhibits significantly more meaningful interactions this can be attributed to the high degree of generalizability of the decomposed language conditioned policy the low level policy generalizing to novel goal images is more robust than generalizing to un grounded language instructions and the high level image generation generalizes well thanks to its internet pre training soar data we also release as a secondary contribution the autonomous dataset collected by our deployment of soar this dataset soar data consists of more than 30 000 trajectories 3m transitions collected with over 50 different sets of objects across 5 different table top setups each trajectory in soar data comes with language annotations from a vlm 5 commanded subgoal images generated by susie during one episode and a task success label predicted by the vlm soar data is similar in size compared to other current robotic datasets but is collected under much smaller time frames in a matter of weeks with minimal human effort in the loop as soar data consists both of failure and success trajectories contains diverse scenes and objects includes language instructions and subgoal images and is collected with a publicly available low cost robotic arm we hope it will be a useful resource for offline reinforcement learning research see https github com rail berkeley soar for instructions on how to download the data bibtex article zhou2024autonomous title autonomous improvement of instruction following skills via foundation models author zhiyuan zhou and pranav atreya and abraham lee and homer walke and oier mees and sergey levine journal arxiv preprint arxiv 407 20635 year 2024 website template borrowed from nerfies under a creative commons attribution sharealike 4 0 international license
|