Meta tags:
Headings (most frequently used words):
-
Text of the page (most frequently used words):
robot (48), the (38), and (35), learning (34), for (30), karl (26), pertsch (26), with (25), conference (19), code (19), that (18), project (17), policies (17), page (16), data (16), paper (16), real (14), arxiv (13), can (12), sergey (12), levine (12), models (12), action (11), our (11), manipulation (11), corl (11), training (11), open (11), dataset (11), from (10), tasks (10), chelsea (10), finn (10), environments (10), based (9), video (9), model (9), task (9), joseph (9), lim (9), approach (9), horizon (9), large (9), policy (9), generalist (9), long (8), skill (8), nbsp (8), website (7), vision (7), learn (7), new (7), world (7), efficient (7), source (7), this (6), performance (6), 2020 (6), scale (6), skills (6), scalable (6), introduce (6), robotics (6), 2024 (6), evaluation (6), evaluations (6), language (6), embodied (6), over (5), propose (5), planning (5), lee (5), reinforcement (5), datasets (5), best (5), imitation (5), meta (5), quan (5), vuong (5), oier (5), mees (5), sim (5), openvla (5), 2025 (5), visual (4), perform (4), control (4), international (4), representations (4), before (4), using (4), simulated (4), systems (4), youngwoon (4), learned (4), via (4), representation (4), human (4), demonstrate (4), embodiment (4), diverse (4), generalization (4), suraj (4), nair (4), kyle (4), vlas (4), reasoning (4), fast (4), roboarena (4), polaris (4), memory (4), pose (3), see (3), iclr (3), kostas (3), daniilidis (3), oleh (3), rybkin (3), what (3), prediction (3), show (3), hierarchical (3), time (3), neurips (3), prior (3), offline (3), oral (3), actions (3), agnostic (3), 2022 (3), approaches (3), collection (3), system (3), enables (3), robots (3), robotic (3), kitchen (3), minutes (3), date (3), trajectories (3), octo (3), supports (3), out (3), full (3), science (3), rss (3), dorsa (3), sadigh (3), droid (3), finalist (3), liang (3), chain (3), thought (3), any (3), william (3), chen (3), develop (3), danny (3), driess (3), short (3), mem (3), google (3), scholar (3), working (3), foundation (3), university (3), phd (3), combining (2), object (2), estimation (2), objects (2), hover (2), image (2), tap (2), screen (2), space (2), predictive (2), poster (2), andrew (2), jaegle (2), kosta (2), derpanis (2), you (2), keyframe (2), keyframes (2), improves (2), zhou (2), future (2), use (2), steps (2), information (2), goal (2), motion (2), enabling (2), them (2), jun (2), yamada (2), transfer (2), when (2), workshop (2), award (2), top (2), demonstrations (2), demonstration (2), quickly (2), common (2), induced (2), parts (2), are (2), enable (2), user (2), teleoperation (2), zhang (2), semantic (2), cross (2), domain (2), between (2), environment (2), three (2), train (2), transformer (2), substantially (2), leads (2), trained (2), range (2), box (2), finetuning (2), release (2), homer (2), walke (2), most (2), collected (2), across (2), scenes (2), improved (2), thomas (2), kollar (2), simpler (2), setups (2), correlation (2), ted (2), xiao (2), evaluating (2), parameter (2), episodes (2), percy (2), moo (2), jin (2), kim (2), like (2), without (2), optimizing (2), mixtures (2), robust (2), weights (2), tokenization (2), method (2), into (2), work (2), brian (2), ichter (2), stachowicz (2), autonomous (2), autoeval (2), pranav (2), atreya (2), strategies (2), distributed (2), benchmark (2), arhan (2), jain (2), correlate (2), strongly (2), key (2), marcel (2), torne (2), was (2), borrowed, layout, here, cnn, regression, dense, surface, labeling, ransac, fitting, accurate, 6dof, texture, less, under, heavy, occlusion, 2018, asian, computer, accv, carsten, rother, eric, brachmann, siva, karthik, mustikovela, omid, hosseini, jafari, ipose, instance, aware, partly, occluded, agent, pure, observations, along, then, used, requiring, orders, magnitude, fewer, annotated, videos, 2019, doing, anything, unsupervisedly, discover, moments, interesting, change, predicted, subgoals, pushing, shenghao, dynamics, jingyun, yang, keyframing, discovery, predicts, sequences, recursive, infilling, devise, allows, mpc, hundreds, neural, processing, dinesh, jayaraman, frederik, ebert, conditioned, predictors, augments, free, agents, capabilities, solve, cluttered, peter, englert, gaurav, sukhatme, max, pflueger, gautam, salhorta, planner, augmented, obstructed, jointly, embedding, tells, guides, effective, which, deep, runner, plenary, talk, accelerating, priors, follow, imitating, demonstrated, instead, primitive, experience, skild, seamlessly, integrate, framework, 2021, yue, guided, extracted, sparse, reward, sung, hwang, shao, hua, sun, taewook, nam, evaluate, effectiveness, visually, complex, substantial, distractors, compare, unsupervised, leverage, scene, important, ignored, anisha, gunjal, assisting, teleoperators, estimates, its, uncertainty, determine, request, input, studies, more, reduced, mental, load, four, parallel, stefanos, nikolaidis, hejia, shivin, dass, assisted, different, even, only, recorded, gopro, strapped, head, akshara, rai, dhruv, batra, franziska, meier, vikash, kumar, ruta, desai, largest, spanning, embodiments, collaboration, 2023, automation, icra, 800k, diffusion, flexible, specification, observation, spaces, configurations, pre, checkpoints, pipelines, tech, report, kevin, black, dibya, ghosh, contains, 76k, 350, hours, interaction, 564, collectors, north, america, asia, europe, course, months, higher, ability, detailed, guide, reproducing, hardware, setup, visualizer, alexander, khazatsky, wild, strong, through, paired, hao, jiajun, jiayuan, hsu, xuanlin, simulation, vla, pretrained, 970k, sets, state, art, controlling, multiple, adapted, fine, tuning, fully, russ, tedrake, benjamin, burchfiel, pannag, sanketi, grace, lam, ethan, foster, rafael, rafailov, ashwin, balakrishna, siddharth, karamcheti, look, think, acting, predict, intermediate, grounded, subtasks, bounding, boxes, etc, increases, challenging, additional, michal, zawalski, group, distributionally, optimization, generates, mixture, outperform, tuned, experts, yichen, jian, chethan, bhateja, joey, hejna, mix, simple, tokenizing, compact, discrete, faster, build, first, zero, shot, finetune, reliable, resets, score, available, public, submit, your, today, tan, zhiyuan, paul, allow, improve, incurring, inference, costs, earlier, ecot, suvir, mirchandani, suneel, belkhale, runs, pairwise, aggregates, global, ranking, wide, providing, comprehensive, tony, pipeline, generating, high, fidelity, clips, results, good, phase, held, also, scores, kanav, arora, abhishek, gupta, mingtong, mixed, modal, combines, text, span, fifteen, cleaning, making, grilled, cheese, sandwich, haohuan, wang, jiaming, tang, karan, dhabalia, 2026, jost, tobias, springenberg, michael, equi, allen, ren, vedder, multi, below, selection, publications, list, interested, machine, moment, towards, focus, challenges, building, developing, scalably, research, linkedin, twitter, email, postdoc, berkeley, stanford, completed, southern, california, usc, during, fortunate, intern, spend, student, researcher, brain, spent, one, year, fulbright, pennsylvania, karol, hausman, member, technical, staff, where, physical, intelligence,
Text of the page (random words):
karl pertsch karl pertsch i am a member of technical staff at physical intelligence where i work on training robot foundation models before that i was a postdoc at uc berkeley and stanford university working with sergey levine and chelsea finn i completed my phd at the university of southern california usc with joseph lim during my phd i was fortunate to intern at meta ai and spend time as a student researcher at google brain with karol hausman before my phd i spent one year as a fulbright scholar at the university of pennsylvania working with kostas daniilidis email nbsp nbsp twitter nbsp nbsp google scholar nbsp nbsp cv nbsp nbsp linkedin research i m interested in machine learning reinforcement learning and robotics at the moment i am working on training foundation models for robotics towards this goal i focus on three key challenges 1 building diverse robot datasets 2 training large scale robot policies on this data and 3 developing approaches for scalably evaluating robot foundation models below is a selection of publications see google scholar for a full list mem multi scale embodied memory for vision language action models marcel torne karl pertsch homer walke kyle vedder suraj nair brian ichter allen z ren haohuan wang jiaming tang kyle stachowicz karan dhabalia michael equi quan vuong jost tobias springenberg sergey levine chelsea finn danny driess arxiv 2026 paper website we introduce mem an approach for mixed modal long horizon memory in robot policies mem combines video based short horizon memory with text based long horizon memory enabling robot policies to perform tasks that span up to fifteen minutes like cleaning up a kitchen or making a grilled cheese sandwich polaris scalable real to sim evaluations for generalist robot policies arhan jain mingtong zhang kanav arora william chen marcel torne abhishek gupta karl pertsch arxiv 2025 paper website code we propose polaris a real to sim pipeline for generating high fidelity simulated environments from short video clips in minutes we demonstrate that evaluation results in polaris environments correlate strongly with real robot performance the key to good real to sim correlation is a short co training phase with sim data collected in held out environments polaris evaluations also correlate strongly with scores in the roboarena benchmark roboarena distributed real world evaluation of generalist robot policies pranav atreya karl pertsch tony lee moo jin kim arhan jain percy liang chelsea finn sergey levine conference on robot learning corl 2025 oral paper website code data we develop roboarena a distributed real world benchmark for generalist robot policies roboarena runs pairwise evaluations on any task in any environment and aggregates them into a global policy ranking this enables scalable robust evaluations across a wide range of tasks and scenes providing the most comprehensive evaluation of generalist robot policies to date training strategies for efficient embodied reasoning william chen suneel belkhale suvir mirchandani oier mees danny driess karl pertsch sergey levine conference on robot learning corl 2025 oral paper website data we propose strategies for efficient embodied reasoning that allow us to improve generalization of robot policies without incurring the inference time costs of our earlier embodied chain of thought ecot approach autoeval autonomous evaluation of generalist robot manipulation policies in the real world zhiyuan paul zhou pranav atreya you liang tan karl pertsch sergey levine conference on robot learning corl 2025 paper website code we develop a system for autonomous evaluation of generalist robot manipulation policies in the real world we finetune models to perform reliable resets and score episodes our system is available to the public submit your policies to autoeval today fast efficient action tokenization for vision language action models karl pertsch kyle stachowicz brian ichter danny driess suraj nair quan vuong oier mees chelsea finn sergey levine robotics science and systems rss 2025 best paper finalist paper website code we release fast a new action tokenization method for vision language action models fast is a simple efficient and scalable method for tokenizing actions into a compact discrete representation with fast we can train vlas 5x faster and build the first vlas that work zero shot in new environments re mix optimizing data mixtures for large scale imitation learning joey hejna chethan bhateja yichen jian karl pertsch dorsa sadigh conference on robot learning corl 2024 best paper finalist paper code we develop a scalable approach for optimizing data mixtures for large scale robot imitation learning using group distributionally robust optimization our approach generates dataset weights for the rt x data mixture that outperform weights tuned by human experts robotic control via embodied chain of thought reasoning michal zawalski william chen karl pertsch oier mees chelsea finn sergey levine conference on robot learning corl 2024 project page paper code models we propose embodied chain of thought learning for vision language action models vlas by training vlas to look and think before acting i e to predict intermediate grounded reasoning steps like subtasks object bounding boxes etc we can enable substantially improved generalization our approach increases the performance of openvla on challenging generalization evaluations by 30 without any additional robot data openvla an open source vision language action model moo jin kim karl pertsch siddharth karamcheti ted xiao ashwin balakrishna suraj nair rafael rafailov ethan foster grace lam pannag sanketi quan vuong thomas kollar benjamin burchfiel russ tedrake dorsa sadigh sergey levine percy liang chelsea finn conference on robot learning corl 2024 best paper finalist project page paper code models we introduce openvla a 7b parameter open source vision language action model vla pretrained on 970k robot episodes from the open x embodiment dataset openvla sets a new state of the art for generalist robot manipulation policies it supports controlling multiple robots out of the box and can be quickly adapted to new robot setups via parameter efficient fine tuning openvla models code and training data are fully open source evaluating real world robot manipulation policies in simulation xuanlin li kyle hsu jiayuan gu karl pertsch oier mees sergey levine jiajun wu chelsea finn hao su quan vuong ted xiao conference on robot learning corl 2024 project page paper code we introduce simpler a collection of simulated environments for manipulation policy evaluation on common real robot setups we demonstrate strong correlation between policy performance in simpler environments and in the real world through paired sim and real evaluations of open source manipulation policies droid a large scale in the wild robot manipulation dataset alexander khazatsky karl pertsch suraj nair thomas kollar sergey levine chelsea finn robotics science and systems rss 2024 project page paper dataset visualizer we introduce droid the most diverse robot manipulation dataset to date it contains 76k demonstration trajectories or 350 hours of interaction data collected across 564 scenes and 84 tasks by 50 data collectors in north america asia and europe over the course of 12 months we demonstrate that training with droid leads to policies with higher performance and improved generalization ability we open source the full dataset policy learning code and a detailed guide for reproducing our robot hardware setup octo an open source generalist robot policy dibya ghosh homer walke karl pertsch kevin black oier mees dorsa sadigh chelsea finn sergey levine robotics science and systems rss 2024 project page tech report code we introduce octo an open source generalist policy trained on 800k robot trajectories octo is a large transformer based diffusion policy that supports flexible task specification observation and action spaces it can control a diverse range of robots out of the box and supports efficient finetuning to new robot configurations we release pre trained checkpoints and our full training finetuning pipelines open x embodiment robotic learning datasets and rt x models open x embodiment collaboration project co leads quan vuong karl pertsch international conference on robotics and automation icra 2023 best conference paper award project page arxiv dataset we introduce the open x embodiment dataset the largest robot learning dataset to date with 1m real robot trajectories spanning 22 robot embodiments we train large transformer based policies on the dataset rt 1 x rt 2 x and show that co training with our diverse dataset substantially improves performance cross domain transfer via semantic skill imitation karl pertsch ruta desai vikash kumar franziska meier joseph j lim dhruv batra akshara rai conference on robot learning corl 2022 project page arxiv code we learn a semantic skill policy that enables cross domain imitation from robot to robot between different environments and even from human video to robot we show that we can learn long horizon robotic manipulation tasks in a simulated kitchen environment using only three minutes of human video recorded in my kitchen with a gopro strapped to my head assisted teleoperation for scalable robot data collection shivin dass karl pertsch hejia zhang youngwoon lee joseph j lim stefanos nikolaidis project page arxiv code we enable scalable robot data collection by assisting human teleoperators with a learned policy our approach estimates its uncertainty over future actions to determine when to request user input in real world user studies we demonstrate that our system enables more efficient teleoperation with reduced mental load and up to four robots in parallel task induced representation learning jun yamada karl pertsch anisha gunjal joseph j lim international conference on learning representations iclr 2022 project page arxiv code we evaluate the effectiveness of representation learning approaches on visually complex environments with substantial distractors we compare common unsupervised representation learning approaches to task induced representations that leverage task information from prior tasks to learn what parts of the scene are important to model and what parts can be ignored skill based meta reinforcement learning taewook nam shao hua sun karl pertsch sung ju hwang joseph j lim international conference on learning representations iclr 2022 project page arxiv code we perform meta rl on top of skills extracted from large task agnostic offline datasets by combining meta training tasks with offline data we can meta learn policies that can quickly learn new long horizon sparse reward tasks demonstration guided reinforcement learning with learned skills karl pertsch youngwoon lee yue wu joseph j lim conference on robot learning corl 2021 project page arxiv code we follow long horizon demonstrations by imitating the demonstrated skills instead of the primitive actions by using skills learned from large task agnostic experience datasets for imitation our approach skild can seamlessly integrate task agnostic data demonstrations via a skill based learning framework accelerating reinforcement learning with learned skill priors karl pertsch youngwoon lee joseph j lim conference on robot learning corl 2020 plenary talk top 4 workshop on robot learning neurips 2020 best paper runner up award deep rl workshop neurips 2020 oral project page arxiv code we jointly learn an embedding space of skills and a prior over skills this skill prior tells us when to use which skill and guides learning on new tasks for effective skill transfer from large offline datasets motion planner augmented reinforcement learning for robot manipulation in obstructed environments jun yamada youngwoon lee gautam salhorta karl pertsch max pflueger gaurav s sukhatme joseph j lim peter englert conference on robot learning corl 2020 project page arxiv code our approach augments model free rl agents with motion planning capabilities enabling them to solve long horizon manipulation tasks in cluttered environments long horizon visual planning with goal conditioned hierarchical predictors karl pertsch oleh rybkin frederik ebert chelsea finn dinesh jayaraman sergey levine conference on neural information processing systems neurips 2020 project page arxiv video code we propose a hierarchical prediction model that predicts sequences by recursive infilling we use this model to devise a hierarchical planning approach that allows to scale visual mpc to long horizon tasks with hundreds of time steps keyframing the future keyframe discovery for visual prediction and planning karl pertsch oleh rybkin jingyun yang shenghao zhou kosta derpanis joseph lim kostas daniilidis andrew jaegle conference on learning for dynamics and control 2020 project page arxiv video poster we propose a keyframe based video prediction model that can unsupervisedly discover the moments of interesting change the keyframes in the data we show that using the predicted keyframes as subgoals for planning improves performance on a simulated pushing task hover over image or tap the screen to see the video learning what you can do before doing anything oleh rybkin karl pertsch kosta derpanis kostas daniilidis andrew jaegle international conference on learning representations iclr 2019 project page arxiv poster we learn an agent s action space from pure visual observations along with a predictive model it can then be used to perform model predictive control requiring orders of magnitude fewer action annotated videos hover over image or tap the screen to see the video ipose instance aware 6d pose estimation of partly occluded objects omid hosseini jafari siva karthik mustikovela karl pertsch eric brachmann carsten rother asian conference on computer vision accv 2018 combining a cnn based regression of dense on object surface labeling with ransac based pose fitting for accurate 6dof pose estimation of texture less objects under heavy occlusion i borrowed this website layout from here
|