Meta tags:
description= Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation;
keywords= Imitation Learning, Cross-Embodiment;
Headings (most frequently used words):
manipulation, crossformer, cross, navigation, locomotion, the, to, of, action, scaling, embodied, learning, for, and, aviation, is, first, robot, policy, achieve, state, art, performance, across, six, embodiments, distinct, spaces, without, any, space, alignment, abstract, approach, results, comparison, best, prior, embodiment, work, bibtex, single, arm, bimanual,
Text of the page (most frequently used words):
the (50), and (32), robot (15), crossformer (13), policy (13), #manipulation (11), #navigation (11), single (11), performance (11), action (10), our (9), embodiment (9), cross (8), learning (8), state (8), art (8), that (7), prior (7), method (7), work (6), arm (6), can (6), embodiments (6), this (5), for (5), alignment (5), across (5), control (5), matches (5), locomotion (4), tasks (4), train (4), first (4), are (4), only (4), any (4), show (4), bimanual (4), into (4), different (4), data (4), each (4), transformer (4), robots (4), from (3), embodied (3), higher (3), complex (3), observation (3), spaces (3), space (3), best (3), successfully (3), level (3), quadrupeds (3), not (3), such (3), goal (3), while (3), without (3), task (3), diverse (3), sequence (3), code (2), scaling (2), aviation (2), ria (2), doshi (2), homer (2), walke (2), oier (2), mees (2), sudeep (2), dasari (2), sergey (2), levine (2), arxiv (2), 2024 (2), sharp (2), turns (2), turn (2), corner (2), obstacle (2), avoidance (2), success (2), section (2), compare (2), demonstrate (2), policies (2), with (2), quadrupedal (2), controls (2), manipulators (2), navigators (2), low (2), most (2), waypoints (2), also (2), capable (2), actions (2), trajectories (2), well (2), more (2), below (2), challenging (2), including (2), specific (2), evaluate (2), which (2), place (2), just (2), trained (2), distinct (2), baseline (2), same (2), architecture (2), better (2), problem (2), based (2), fed (2), dataset (2), systems (2), model (2), datasets (2), generalization (2), have (2), training (2), website, was, adapted, following, source, article, doshi24, title, one, author, journal, preprint, 2408, 11812, year, bibtex, forgoes, achieves, average, rates, video, rollouts, who, aligning, though, positive, transfer, find, sometimes, hinders, third, person, qualitatively, smoother, less, stop, movement, yang, comparison, must, learn, walk, forward, motion, related, using, high, but, handling, dimensional, needed, entirely, given, topological, map, goals, should, follow, path, avoiding, collisions, navigating, both, indoors, outdoors, out, distribution, paths, addition, strong, turning, see, uncap, orange, sharpie, pen, successful, due, various, specifications, longer, horizon, chunking, frequency, 50hz, flexibility, allows, meet, constraints, adversely, impacting, other, franka, sweeping, grasp, grey, brush, sweep, acorns, dustpan, widowx, two, pick, put, mushroom, silver, pot, spoon, blue, towel, possible, variety, limited, maintaining, achieved, provide, evaluation, videos, four, modes, here, refers, whichever, has, rate, target, overall, performs, than, comparably, results, casts, imitation, choose, handle, variable, length, inputs, outputs, observations, proprioception, specification, tokenized, modality, tokenizers, these, assembled, token, decoder, backbone, shared, output, embeddings, then, separate, heads, class, produce, corresponding, dimension, all, approach, propose, scalable, flexible, consume, largest, date, 900k, network, weights, vastly, dual, wheeled, quadcopters, unlike, does, require, manual, extensive, experiments, real, world, specialist, tailored, significantly, outperforming, modern, machine, rely, large, attain, broad, often, poses, challenge, where, robotic, platform, might, small, many, kinds, leverage, much, broader, lead, robustness, however, multi, because, widely, varying, sensors, actuators, frequencies, abstract, achieve, six, paper, equal, contribution, berkeley, carnegie, mellon, university, conference, corl, oral, presentation, top,
Text of the page (random words):
crossformer crossformer scaling cross embodied learning for manipulation navigation locomotion and aviation oral presentation top 4 conference on robot learning corl 2024 ria doshi 1 homer walke 1 oier mees 1 sudeep dasari 2 sergey levine 1 equal contribution 1 uc berkeley 2 carnegie mellon university paper code model crossformer is the first robot policy to achieve state of the art performance across six embodiments of distinct action spaces without any action space alignment abstract modern machine learning systems rely on large datasets to attain broad generalization and this often poses a challenge in robot learning where each robotic platform and task might have only a small dataset by training a single policy across many different kinds of robots a robot learning method can leverage much broader and more diverse datasets which in turn can lead to better generalization and robustness however training a single policy on multi robot data is challenging because robots can have widely varying sensors actuators and control frequencies we propose crossformer a scalable and flexible transformer based policy that can consume data from any embodiment we train crossformer on the largest and most diverse dataset to date 900k trajectories across 30 different robot embodiments we demonstrate that the same network weights can control vastly different robots including single and dual arm manipulation systems wheeled robots quadcopters and quadrupeds unlike prior work our model does not require manual alignment of the observation or action spaces extensive experiments in the real world show that our method matches the performance of specialist policies tailored for each embodiment while also significantly outperforming the prior state of the art in cross embodiment learning approach crossformer architecture our method crossformer casts the cross embodied imitation learning problem as a sequence to sequence problem we choose a transformer based policy to handle the variable length inputs and outputs observations proprioception and a task specification are tokenized by modality specific tokenizers these are assembled into a token sequence and fed into a decoder only transformer backbone that is shared across all embodiments the output embeddings of this transformer are then fed into separate action heads for each class of embodiments to produce actions of the corresponding dimension results we compare crossformer to the same architecture trained on just the target embodiment data single robot baseline as well as the best prior method overall crossformer performs better than or comparably to the state of the art below we provide evaluation videos for each of our four distinct action modes single arm manipulation bimanual manipulation navigation and locomotion here the state of the art method refers to the best prior method or the single robot baseline whichever has a higher success rate single arm manipulation on the franka we evaluate on a sweeping task in which the robot s goal is to grasp the grey brush and sweep the acorns into the dustpan on the widowx we evaluate on two different pick and place tasks put the mushroom in the silver pot and place the spoon on the blue towel matches state of the art performance ︎ we show that it is possible to co train a policy on a variety of diverse embodiments not just limited to single arm manipulators while maintaining performance achieved by prior work trained only on single arm data bimanual manipulation the robot s goal is to uncap the orange sharpie pen matches state of the art performance ︎ crossformer is first work to show successful performance of a cross embodiment policy on a bimanual robot the bimanual embodiment is challenging due to various specifications including longer horizon action chunking and a higher control frequency 50hz our method s flexibility allows us to meet such robot specific constraints without adversely impacting performance on other embodiments navigation given a topological map of goals the navigation robot should follow the path to the goal while avoiding any collisions matches state of the art performance ︎ without any action space alignment our policy is capable of navigating trajectories both indoors and outdoors as well as on out of distribution paths in addition we show strong performance on more complex navigation tasks such as obstacle avoidance sharp turns and corner turning see the section below locomotion the quadrupedal robot must learn to successfully walk forward matches state of the art performance ︎ our policy is the first to successfully train a cross embodiment policy that successfully controls manipulators navigators and low level motion of quadrupeds the most related cross embodiment prior work controls quadrupeds using high level 2d waypoints crossformer can not only control such navigators with 2d waypoints but is also capable of handling the low level 12 dimensional actions needed to entirely control a quadrupedal robot comparison to best prior cross embodiment work in this section we compare video rollouts of our policy to yang et al who train a single policy across manipulation and navigation by aligning the observation and action spaces though this work is the first to demonstrate positive transfer from navigation to manipulation we find that observation and action space alignment sometimes hinders performance on complex navigation and third person single arm manipulation tasks qualitatively our navigation policies are smoother with less stop go movement crossformer forgoes action alignment and achieves an average of 3x higher success rates on complex navigation and manipulation tasks obstacle avoidance turn corner sharp turns bibtex article doshi24 crossformer title scaling cross embodied learning one policy for manipulation navigation locomotion and aviation author ria doshi and homer walke and oier mees and sudeep dasari and sergey levine journal arxiv preprint arxiv 2408 11812 year 2024 this website was adapted from the following source code
|