Meta tags:
description= PaLM-E: An Embodied Multimodal Language Model;
Headings (most frequently used words):
the, blocks, to, push, demo, bring, me, green, abstract, approach, results, citation, acknowledgements, paper, rice, chips, from, drawer, star, sort, by, colors, into, different, corners, incorporating, visual, feedback, over, long, time, horizons, move, remaining, group, ocean, colored, together, red, coffee, cup, turtle,
Text of the page (most frequently used words):
the (52), and (38), language (20), palm (16), model (15), #blocks (11), embodied (9), visual (8), robot (7), example (6), from (6), our (6), green (6), with (6), show (6), long (6), into (6), tasks (6), that (6), push (5), instruction (5), multiple (5), trained (5), continuous (5), for (4), multimodal (4), planning (4), next (4), another (4), this (4), move (4), over (4), horizon (4), star (4), bring (4), chowdhery (3), arxiv (3), prompt (3), below (3), all (3), previous (3), turtle (3), red (3), coffee (3), cup (3), successfully (3), execute (3), group (3), incorporating (3), feedback (3), different (3), can (3), plan (3), input (3), real (3), same (3), state (3), modalities (3), pre (3), large (3), models (3), driess (2), danny (2), xia (2), fei (2), sajjadi (2), mehdi (2), lynch (2), corey (2), aakanksha (2), ichter (2), brian (2), wahid (2), ayzaan (2), tompson (2), jonathan (2), vuong (2), quan (2), tianhe (2), huang (2), wenlong (2), chebotar (2), yevgen (2), sermanet (2), pierre (2), duckworth (2), daniel (2), levine (2), sergey (2), vanhoucke (2), vincent (2), hausman (2), karol (2), toussaint (2), marc (2), greff (2), klaus (2), zeng (2), andy (2), mordatch (2), igor (2), florence (2), pete (2), orange (2), text (2), gray (2), examples (2), are (2), completions (2), more (2), images (2), demo (2), addition (2), capabilities (2), vision (2), please (2), paper (2), demonstrate (2), two (2), generalization (2), only (2), them (2), where (2), able (2), task (2), remaining (2), time (2), horizons (2), sort (2), colors (2), corners (2), stages (2), finally (2), first (2), step (2), such (2), rice (2), chips (2), drawer (2), embodiments (2), these (2), results (2), directly (2), observations (2), sensor (2), embedding (2), space (2), tokens (2), textual (2), world (2), robotics (2), encodings (2), end (2), variety (2), scale (2), 562b (2), generalist (2), authors, would, like, thank, their, advice, help, support, chen, etienne, pot, sebastian, goodman, ted, xiao, keerthana, gopalakrishnan, kehang, han, henryk, michalewski, neil, houlsby, basil, mustafa, justin, gilmer, yonghui, erica, moreira, victor, gomes, tom, duerig, kendra, byrne, acknowledgements, inproceedings, driess2023palme, title, author, booktitle, preprint, 2303, 03378, year, 2023, version, citation, response, shade, one, unlocking, new, competent, check, out, details, see, dmeo, case, dataset, contains, three, demonstrations, none, included, even, though, has, never, seen, before, ocean, colored, together, following, part, controlling, table, top, arranging, based, pushing, sequences, commands, low, level, policy, yellow, hexagon, blue, triangle, few, videos, showing, how, used, note, were, obtained, using, data, video, includes, steps, well, camera, object, wasn, exposed, main, architectural, idea, inject, estimates, other, realized, encoding, sequence, vectors, dimension, information, hence, injected, analogous, way, decoder, llm, generates, autoregressively, given, prefix, call, since, use, make, mbodied, 2022, approach, have, been, demonstrated, perform, complex, however, enabling, general, inference, problems, raises, challenge, grounding, propose, incorporate, thereby, establish, link, between, words, percepts, multi, modal, sentences, interleave, estimation, train, conjunction, including, sequential, robotic, manipulation, question, answering, captioning, evaluations, single, address, reasoning, observation, further, exhibits, benefits, diverse, joint, training, across, internet, domains, largest, parameters, being, art, performance, vqa, retains, increasing, positive, transfer, abstract,
Text of the page (random words):
palm e an embodied multimodal language model palm e an embodied multimodal language model danny driess 1 2 fei xia 1 mehdi s m sajjadi 3 corey lynch 1 aakanksha chowdhery 3 brian ichter 1 ayzaan wahid 1 jonathan tompson 1 quan vuong 1 tianhe yu 1 wenlong huang 1 yevgen chebotar 1 pierre sermanet 1 daniel duckworth 3 sergey levine 1 vincent vanhoucke 1 karol hausman 1 marc toussaint 2 klaus greff 3 andy zeng 1 igor mordatch 3 pete florence 1 1 2 3 paper demo abstract large language models have been demonstrated to perform complex tasks however enabling general inference in the real world e g for robotics problems raises the challenge of grounding we propose embodied language models to directly incorporate real world continuous sensor modalities into language models and thereby establish the link between words and percepts input to our embodied language model are multi modal sentences that interleave visual continuous state estimation and textual input encodings we train these encodings end to end in conjunction with a pre trained large language model for multiple embodied tasks including sequential robotic manipulation planning visual question answering and captioning our evaluations show that palm e a single large embodied multimodal model can address a variety of embodied reasoning tasks from a variety of observation modalities on multiple embodiments and further exhibits positive transfer the model benefits from diverse joint training across internet scale language vision and visual language domains our largest model palm e 562b with 562b parameters in addition to being trained on robotics tasks is a visual language generalist with state of the art performance on ok vqa and retains generalist language capabilities with increasing scale approach the main architectural idea of palm e is to inject continuous embodied observations such as images state estimates or other sensor modalities into the language embedding space of a pre trained language model this is realized by encoding the continuous observations into a sequence of vectors with the same dimension as the embedding space of the language tokens the continuous information is hence injected into the language model in an analogous way to language tokens palm e is a decoder only llm that generates textual completions autoregressively given a prefix or prompt we call our model palm e since we use palm chowdhery et al 2022 as the pre trained language model and make it e mbodied results we show a few example videos showing how palm e can be used to plan and execute long horizon tasks on two different real embodiments please note that all of these results were obtained using the same model trained on all data in the first video we execute a long horizon instruction bring me the rice chips from the drawer that includes multiple planning steps as well as incorporating visual feedback from the robot s camera finally show another example on the same robot where the instruction is bring me a green star green star is an object that this robot wasn t directly exposed to bring me the rice chips from the drawer bring me the green star previous next in the following part we show palm e controlling a table top robot arranging blocks we show the palm e can successfully plan over multiple stages based on visual and language input our model is able to successfully plan a long horizon task sort blocks by colors into different corners another example of planning over multiple stages and incorporating visual feedback over long time horizons finally we demonstrate another example of long horizon pushing tasks on this robot the first instruction is move remaining blocks to the group palm e sequences step by step commands to the low level policy such as move the yellow hexagon to the green star and move the blue triangle to the group sort blocks by colors into different corners incorporating visual feedback over long time horizons move remaining blocks to the group push the ocean colored blocks together previous next next we demonstrate two examples of generalization in the case below the instruction is push red blocks to the coffee cup the dataset contains only three demonstrations with the coffee cup in them and none of them included red blocks we show another generalization example where the instruction is push green blocks to the turtle the robot is able to successfully execute this task even though it has never seen the turtle before push red blocks to the coffee cup push green blocks to the turtle previous next in addition to unlocking new capabilities in robot planning palm e is a competent vision language model please check out our paper for more details and see the dmeo below demo the examples below are all example completions in orange from palm e the prompt is the one or more images and the text in gray prompt text in gray palm e response in orange shade citation arxiv version inproceedings driess2023palme title palm e an embodied multimodal language model author driess danny and xia fei and sajjadi mehdi s m and lynch corey and chowdhery aakanksha and ichter brian and wahid ayzaan and tompson jonathan and vuong quan and yu tianhe and huang wenlong and chebotar yevgen and sermanet pierre and duckworth daniel and levine sergey and vanhoucke vincent and hausman karol and toussaint marc and greff klaus and zeng andy and mordatch igor and florence pete booktitle arxiv preprint arxiv 2303 03378 year 2023 acknowledgements the authors would like to thank for their advice help and support xi chen etienne pot sebastian goodman ted xiao keerthana gopalakrishnan kehang han henryk michalewski neil houlsby basil mustafa justin gilmer yonghui wu erica moreira victor gomes tom duerig and kendra byrne
|