Meta tags:
description= ;
author= ;
Headings (most frequently used words):
-
Text of the page (most frequently used words):
the (62), learning (56), and (53), link (40), human (38), from (30), with (29), for (28), feedback (28), can (21), that (21), interactive (20), how (16), implicit (15), this (15), university (14), reward (13), humans (11), learn (11), such (11), language (10), will (10), are (10), research (9), reinforcement (9), task (9), models (9), talk (9), systems (9), which (8), about (8), abstract (8), policy (7), using (7), interaction (7), rewards (7), contextual (7), agent (7), user (7), their (7), when (7), some (7), our (7), training (6), via (6), yang (6), based (6), bandits (6), design (6), spotlight (6), facial (6), signals (6), workshop (6), microsoft (5), information (5), wang (5), david (5), robot (5), data (5), session (5), real (5), large (5), these (5), first (5), discuss (5), during (5), leverage (5), enable (5), questions (4), john (4), organizers (4), active (4), offline (4), control (4), through (4), environment (4), cognitive (4), natural (4), under (4), wen (4), sun (4), huang (4), mapping (4), space (4), also (4), abel (4), taylor (4), kessler (4), faulkner (4), bradley (4), knox (4), paul (4), mineiro (4), robots (4), provide (4), algorithms (4), work (4), need (4), grounded (4), dogmas (4), what (4), agents (4), reactions (4), other (4), could (4), jesse (4), thomason (4), possible (4), should (4), langford (3), california (3), berkeley (3), anca (3), dragan (3), mengdi (3), jonathan (3), preference (3), dorsa (3), sadigh (3), driven (3), ayush (3), imitation (3), karthik (3), wenhao (3), generative (3), steven (3), demonstrations (3), kim (3), decision (3), into (3), preferences (3), sam (3), furong (3), haifeng (3), interface (3), daniel (3), optimal (3), latent (3), fan (3), eye (3), prediction (3), papers (3), keerthana (3), gopalakrishnan (3), improve (3), teachers (3), people (3), however (3), imperfect (3), address (3), have (3), has (3), robotics (3), washington (3), break (3), poster (3), recent (3), three (3), all (3), well (3), gestures (3), expressions (3), understanding (3), guidance (3), stage (3), ability (3), problems (3), hope (3), initially (3), unknown (3), even (3), embodied (3), applications (3), pattern (3), google (3), deepmind (3), non (3), machine (3), better (3), itself (3), challenges (3), sponsors (2), out (2), pre (2), choices (2), conditional (2), inverse (2), without (2), function (2), towards (2), optimality (2), fine (2), tuned (2), diverse (2), improving (2), calibrate (2), community (2), set (2), model (2), misspecification (2), shuo (2), sujay (2), sanghavi (2), bayesian (2), sabina (2), sekhari (2), sridharan (2), runzhe (2), zhan (2), masatoshi (2), uehara (2), jason (2), lee (2), lin (2), ren (2), types (2), specific (2), finale (2), doshi (2), velez (2), harris (2), shared (2), multi (2), gokul (2), swamy (2), ding (2), zhao (2), sanjiban (2), choudhury (2), online (2), thomas (2), relative (2), optimization (2), image (2), settings (2), post (2), text (2), rlhf (2), remarks (2), panel (2), deployed (2), performance (2), input (2), ways (2), often (2), current (2), lab (2), lens (2), working (2), coffee (2), contributed (2), but (2), scalar (2), limits (2), where (2), must (2), assumptions (2), conclude (2), been (2), shaped (2), rather (2), than (2), second (2), solution (2), argue (2), two (2), while (2), interactions (2), its (2), teaching (2), methods (2), problem (2), then (2), propose (2), novel (2), framework (2), empathic (2), live (2), domain (2), manipulation (2), bowman (2), jim (2), diyi (2), brown (2), suboptimal (2), ambiguous (2), pitfalls (2), overcome (2), use (2), grizou (2), matrix (2), 4th (2), attempt (2), whose (2), meaning (2), internal (2), ptlm (2), researchers (2), usually (2), next (2), word (2), actions (2), loop (2), multimodal (2), those (2), context (2), role (2), rich (2), aligned (2), allow (2), machines (2), time (2), schedule (2), texas (2), austin (2), stanford (2), usc (2), amazon (2), panelists (2), speakers (2), below (2), arbitrary (2), hci (2), assist (2), build (2), much (2), account (2), changes (2), speech (2), movements (2), etc (2), beyond (2), grounding (2), end (2), users (2), world (2), tagged (2), intent (2), may (2), reach, any, gmail, com, contact, inria, pierre, yves, oduyer, illinois, urbana, champaign, tengyang, xie, andreea, bobu, massachusetts, institute, technology, andi, peng, akanksha, saran, guided, search, parameterized, skills, adverbs, benjamin, adin, spiegel, george, konidaris, crowd, sourcing, improves, retrieval, zhuotong, chen, yifei, branislav, kveton, anoop, deoras, accelerating, exploration, representation, bogdan, mazoure, jake, bruce, doina, precup, rob, fergus, ankit, anand, dynamic, pessimism, zihao, zhuoran, chain, prompting, ambiguity, resolution, crosslingual, generation, pilault, xavier, garcia, arthur, brazinskas, orhan, firat, joey, hejna, rewarded, soups, pareto, interpolating, weights, alexandre, rame, guillaume, couairon, corentin, dancette, mustafa, shukor, jean, baptiste, gaya, laure, soulier, matthieu, cord, bionic, limb, batch, game, kilian, freitag, rita, laezza, jan, zbinden, max, ortiz, catalan, modeled, uncertainty, jaelle, scheuerman, zachary, bishof, chris, michael, building, libraries, programs, leonardo, hernandez, cano, yewen, robert, hawkins, joshua, tenenbaum, armando, solar, lezama, selection, rajat, sen, meta, prior, sloman, bharti, samuel, kaski, queries, query, efficiently, swiftsage, fast, slow, thinking, complex, tasks, bill, yuchen, yicheng, karina, prithviraj, ammanabrolu, faeze, brahman, shiyu, chandra, bhagavatula, yejin, choi, xiang, provable, nathan, kallus, discovering, traits, behaviors, lars, lien, ankile, brian, ham, kevin, mao, eura, shin, siddharth, swaroop, weiwei, pan, strategic, apple, tasting, keegan, chara, podimata, safety, constraints, konwoo, zuxin, liu, selective, sampling, regression, unraveling, arc, puzzle, mimicking, solutions, object, centric, transformer, jaehyun, park, jaegyun, sanha, hwang, mintaek, lim, ualibekova, sejin, sundong, simulators, tap, ardavan, nobandegani, shultz, irina, rish, complementing, different, observation, drew, bagnell, behavioral, attributes, filling, gap, between, symbolic, goal, specification, guan, valmeekam, subbarao, kambhampati, temporally, extended, prompts, medical, segmentation, chuyun, shen, zhang, xiangfeng, principal, alignment, bilevel, souradip, chakraborty, amrit, bedi, alec, koppel, transition, leo, benac, sonali, parbhoo, follow, ups, matter, serving, contexts, chaoqi, ziyu, zhe, feng, ashwinkumar, badanidiyuru, minecraft, shalev, lifshitz, keiran, paster, chan, jimmy, sheila, mcilraith, meet, mechanism, combat, clickbait, recommendation, kleine, buening, aadirupa, saha, christos, dimitrakakis, blender, configurable, yannick, metz, lindner, raphaël, baur, keim, mennatallah, assady, asymptotically, fixed, budget, best, arm, identification, variance, dependent, bounds, masahiro, kato, masaaki, imaizumi, takuya, ishihara, toru, kitagawa, legible, motion, matthew, bronars, danfei, visual, encoding, jielin, qiu, william, han, ucb, provably, learns, inconsistent, tongzheng, inderjit, dhillon, survival, instinct, bias, anqi, dipendra, misra, andrey, kolobov, ching, cheng, recommendations, yao, chuanhao, denis, nekipelov, hongning, gaze, objective, ravi, kumar, thakur, sunbeam, vinicius, goecks, ellen, novoseller, ritwik, bera, vernon, lawhern, greg, gremillion, valasek, nicholas, waytowich, concluding, wild, furthermore, both, benefit, adapt, around, them, act, unable, quantities, past, conducted, toward, addressing, issues, focused, creating, assisted, feeding, project, personal, approaching, similar, possibly, talks, highly, practical, specify, adoption, motivates, study, inferred, observables, aka, theoretic, argument, indicates, additional, succeed, review, sufficient, conditions, literature, speculation, composing, igl, modern, part, call, refers, focus, environments, treatment, finding, endless, adaptation, last, hypothesis, states, goals, purposes, thought, maximization, signal, views, ought, dispense, entirely, recognize, embrace, nuance, third, vocalizations, abundant, naturally, occurring, channel, cost, approach, contrasts, common, critiques, attentively, intentionally, provided, define, general, method, consists, relevant, statistics, advantage, instantiate, evaluations, learned, collect, dataset, participants, observe, execute, sub, prescribed, train, deep, neural, network, demonstrate, infer, ranking, events, prerecorded, transfer, evaluates, trajectories, lunch, incomplete, biased, leading, misidentification, true, behavior, seeks, techniques, biases, multiple, align, feature, representations, interpretable, paths, forward, 2x2, position, corner, yet, fully, explored, efforts, positioning, highlight, direction, expose, barriers, effort, demonstration, iftt, pin, self, calibrating, developed, permits, aiming, consistency, pillar, pretrained, rage, right, now, perspective, folks, who, intersection, vision, since, before, was, cool, noticeable, impact, outside, nlp, feel, like, they, plug, exclusively, trained, only, potentially, words, thousands, underpaid, annotators, means, involved, images, captions, describe, literal, content, extract, kinds, cover, worked, broader, open, appropriate, ptlms, considering, instructions, along, autonomy, robotic, creative, tapping, more, specifically, few, vignettes, llms, vlms, social, reasoning, corrective, finally, discussing, effective, identify, patterns, token, invariant, fashion, transformation, extrapolation, show, evidence, solving, era, introductory, gmt, maryland, nvidia, nyu, anthropic, glasgow, utah, archival, submitted, venues, ask, embedded, page, sli, potential, listed, minimal, paradigm, known, imported, massively, used, missing, today, technological, paradigms, scale, accessibility, communities, adaptive, interfaces, targeting, wide, range, marginalized, specially, abled, sections, society, intrinsic, push, become, socially, integrated, coordinated, develop, interact, teach, lead, designs, average, versus, personalized, finetuning, stationary, over, stationarity, meanings, there, explicit, external, hand, crafted, high, dimensional, interactively, quickly, becoming, widespread, typically, offer, wealth, cues, form, process, create, thereby, closed, sequential, making, offers, unique, distribution, influenced, algorithm, thus, adaptively, nature, rapidly, express, various, forms, amenable, naturalistic, going, organizing, bring, together, interdisciplinary, experts, computer, science, explore, foster, discussions, envision, exchange, ideas, within, across, disciplines, new, bridges, most, valuable, young, interested, growing, careers, check, recording, gather, town, virtual, location, hawaii, convention, center, room, 315, icml, 2023, saturday, july, 29th,
Text of the page (random words):
ion that this exchange of ideas within and across disciplines can build new bridges address some of the most valuable challenges in interactive learning with implicit human feedback and also provide guidance to young researchers interested in growing their careers in this space some potential questions we hope to discuss at this workshop are listed below when is it possible to go beyond reinforcement learning with hand crafted rewards and leverage interaction grounded learning from arbitrary feedback signals where grounding for such feedback could be initially unknown contextual rich and high dimensional how can we learn from natural implicit human feedback signals such as natural language speech eye movements facial expressions gestures etc during interaction is it possible to learn from human guidance signals whose meanings are initially unknown or ambiguous even when there is no explicit external reward how should learning algorithms account for a human s preferences or internal reward that is non stationary and changes over time how can we account for non stationarity of the environment itself how much of the learning should be pre training i e learning for the average user versus how much should it be interactive or personalized i e for finetuning to a specific user how can we develop a better understanding of how humans interact with teach other humans or machines and how could such an understanding lead to better designs for learning systems that leverage human signals during interaction how to design intrinsic reward systems that could push agents to learn to become socially integrated coordinated aligned with humans how can well known design methods from hci such as ability based design be imported and massively used in ai ml what is missing from today s technological solution paradigms that can allow for ability based design to be deployed at scale how can the machine learning community assist hci and accessibility research communities to build adaptive learning interfaces targeting a wide range of marginalized and specially abled sections of society what are the minimal set of assumptions under which learning from arbitrary implicit feedback signals is possible for the interaction grounded learning paradigm all our contributed papers are non archival and can be submitted to other venues to ask questions during the workshop use this sli do link or the embedded page below speakers david abel google deepmind daniel brown university of utah jonathan grizou university of glasgow taylor kessler faulkner university of washington bradley knox university of texas austin paul mineiro microsoft research dorsa sadigh stanford university jesse thomason usc and amazon panelists sam bowman nyu and anthropic jim fan nvidia furong huang university of maryland jesse thomason usc and amazon diyi yang stanford university david abel google deepmind anca dragan university of california berkeley keerthana gopalakrishnan google deepmind taylor kessler faulkner university of washington bradley knox university of texas austin john langford microsoft research paul mineiro microsoft research schedule time gmt 10 09 00 am 09 10 am organizers introductory remarks 09 10 am 09 35 am dorsa sadigh interactive learning in the era of large models abstract in this talk i will discuss the role of language in learning from interactions with humans i will first talk about how language instructions along with latent actions can enable shared autonomy in robotic manipulation problems i will then talk about creative ways of tapping into the rich context of large models to enable more aligned ai agents specifically i will discuss a few vignettes about how we can leverage llms and vlms to learn human preferences allow for grounded social reasoning or enable teaching humans using corrective feedback i will finally conclude the talk by discussing how large models can be effective pattern machines that can identify patterns in a token invariant fashion and enable pattern transformation extrapolation and even show some evidence of pattern optimization for solving control problems 09 35 am 10 00 am jesse thomason considering the role of language in embodied systems abstract pretrained language models ptlm are all the rage right now from the perspective of folks who have been working at the intersection of language vision and robotics since before it was cool the noticeable impact is that researchers outside nlp feel like they should plug language into their work however these models are exclusively trained on text data usually only for next word prediction and potentially for next word prediction but under a fine tuned words as actions policy with thousands of underpaid human annotators in the loop e g rlhf even when a ptlm is multimodal that usually means training also involved images and their captions which describe the literal content of the image what meaning can we hope to extract from those kinds of models in the context of embodied interactive systems in this talk i ll cover some applications our lab has worked through in the space language and embodied systems with a broader lens towards open questions about the limits and in appropriate applications of current ptlms with those systems 10 00 am 10 30 am coffee break poster session 10 30 am 10 55 am jonathan grizou aiming for internal consistency the 4th pillar of interactive learning abstract i will propose a 2x2 matrix to position interactive learning systems and argue that the 4th corner of that space is yet to be fully explored by our research efforts by positioning recent work on that matrix i hope to highlight a possible research direction and expose barriers to be overcome in that effort i will attempt a live demonstration of iftt pin a self calibrating interface we developed that permits a user to control an interface using signals whose meaning are initially unknown 10 55 am 11 20 am daniel brown pitfalls and paths forward when learning rewards from human feedback abstract human feedback is often incomplete suboptimal biased and ambiguous leading to misidentification of the human s true reward function and suboptimal agent behavior i will discuss these pitfalls as well as some of our recent work that seeks to overcome these problems via techniques that calibrate to user biases learn from multiple feedback types use human feedback to align robot feature representations and enable interpretable reward learning 11 20 am 12 10 pm panel session 1 sam bowman jim fan furong huang jesse thomason diyi yang keerthana gopalakrishnan 12 10 pm 01 10 pm lunch break 01 10 pm 01 35 pm bradley knox the empathic framework for task learning from implicit human feedback abstract reactions such as gestures facial expressions and vocalizations are an abundant naturally occurring channel of information that humans provide during interactions a robot or other agent could leverage an understanding of such implicit human feedback to improve its task performance at no cost to the human this approach contrasts with common agent teaching methods based on demonstrations critiques or other guidance that need to be attentively and intentionally provided in this talk we first define the general problem of learning from implicit human feedback and then propose to address this problem through a novel data driven framework empathic this two stage method consists of 1 mapping implicit human feedback to relevant task statistics such as rewards optimality and advantage and 2 using such a mapping to learn a task we instantiate the first stage and three second stage evaluations of the learned mapping to do so we collect a dataset of human facial reactions while participants observe an agent execute a sub optimal policy for a prescribed training task we train a deep neural network on this data and demonstrate its ability to 1 infer relative reward ranking of events in the training task from prerecorded human facial reactions 2 improve the policy of an agent in the training task using live human facial reactions and 3 transfer to a novel domain in which it evaluates robot manipulation trajectories 01 35 pm 02 00 pm david abel three dogmas of reinforcement learning abstract modern reinforcement learning has been in large part shaped by three dogmas the first is what i call the environment spotlight which refers to our focus on environments rather than agents the second is our implicit treatment of learning as finding a solution rather than endless adaptation the last is the reward hypothesis which states that all goals and purposes can be well thought of as maximization of a reward signal in this talk i discuss how these dogmas have shaped our views on learning i argue that when agents learn from human feedback we ought to dispense entirely with the first two dogmas while we must recognize and embrace the nuance implicit in the third 02 00 pm 02 25 pm paul mineiro contextual bandits without rewards abstract contextual bandits are highly practical but the need to specify a scalar reward limits their adoption this motivates study of contextual bandits where a latent reward must be inferred from post decision observables aka interactive grounded learning an information theoretic argument indicates the need for additional assumptions to succeed and i review some sufficient conditions from the recent literature i conclude with speculation about composing igl with active learning 02 25 pm 03 00 pm contributed talks 03 00 pm 03 30 pm coffee break poster session 03 30 pm 04 00 pm taylor kessler faulkner robots learning from real people abstract robots deployed in the wild can improve their performance by using input from human teachers furthermore both robots and humans can benefit when robots adapt to and learn from the people around them however real people can act in imperfect ways and can often be unable to provide input in large quantities in this talk i will address some of the past research i have conducted toward addressing these issues which has focused on creating learning algorithms that can learn from imperfect teachers i will also talk about my current work on the robot assisted feeding project in the personal robotics lab at the university of washington which i am approaching through a similar lens of working with real teachers and possibly imperfect information 04 00 pm 04 50 pm panel session 2 david abel anca dragan keerthana gopalakrishnan taylor kessler faulkner bradley knox john langford paul mineiro 04 50 pm 05 00 pm organizers concluding remarks papers imitation learning with human eye gaze via multi objective prediction link spotlight ravi kumar thakur md sunbeam vinicius g goecks ellen novoseller ritwik bera vernon lawhern greg gremillion john valasek nicholas r waytowich learning from a learning user for optimal recommendations link spotlight fan yao chuanhao li denis nekipelov hongning wang haifeng xu survival instinct in offline reinforcement learning and implicit human bias in data link spotlight anqi li dipendra misra andrey kolobov ching an cheng ucb provably learns from inconsistent human feedback link spotlight shuo yang tongzheng ren inderjit s dhillon sujay sanghavi visual based policy learning with latent language encoding link spotlight jielin qiu mengdi xu william han bo li ding zhao legible robot motion from conditional generative models link matthew bronars danfei xu asymptotically optimal fixed budget best arm identification with variance dependent bounds link masahiro kato masaaki imaizumi takuya ishihara toru kitagawa rlhf blender a configurable interactive interface for learning from diverse human feedback link yannick metz david lindner raphaël baur daniel a keim mennatallah el assady bandits meet mechanism design to combat clickbait in online recommendation link thomas kleine buening aadirupa saha christos dimitrakakis haifeng xu a generative model for text control in minecraft link shalev lifshitz keiran paster harris chan jimmy ba sheila a mcilraith follow ups also matter improving contextual bandits via post serving contexts link chaoqi wang ziyu ye zhe feng ashwinkumar badanidiyuru haifeng xu bayesian inverse transition learning for offline settings link leo benac sonali parbhoo finale doshi velez principal driven reward design and agent policy alignment via bilevel rl link souradip chakraborty amrit bedi alec koppel furong huang mengdi wang temporally extended prompts optimization for sam in interactive medical image segmentation link chuyun shen wenhao li ya zhang xiangfeng wang relative behavioral attributes filling the gap between symbolic goal specification and reward learning from human preferences link lin guan karthik valmeekam subbarao kambhampati complementing a policy with a different observation space link gokul swamy sanjiban choudhury drew bagnell steven wu cognitive models as simulators using cognitive models to tap into implicit human feedback link ardavan s nobandegani thomas shultz irina rish unraveling the arc puzzle mimicking human solutions with object centric decision transformer link jaehyun park jaegyun im sanha hwang mintaek lim sabina ualibekova sejin kim sundong kim selective sampling and imitation learning via online regression link ayush sekhari karthik sridharan wen sun runzhe wu learning shared safety constraints from multi task demonstrations link konwoo kim gokul swamy zuxin liu ding zhao sanjiban choudhury steven wu strategic apple tasting link keegan harris chara podimata steven wu discovering user types mapping user traits by task specific behaviors in reinforcement learning link lars lien ankile brian ham kevin mao eura shin siddharth swaroop finale doshi velez weiwei pan provable offline reinforcement learning with human feedback link wenhao zhan masatoshi uehara nathan kallus jason d lee wen sun swiftsage a generative agent with fast and slow thinking for complex interactive tasks link bill yuchen lin yicheng fu karina yang prithviraj ammanabrolu faeze brahman shiyu huang chandra bhagavatula yejin choi xiang ren how to query human feedback efficiently in rl link wenhao zhan masatoshi uehara wen sun jason d lee contextual bandits and imitation learning with preference based active queries link ayush sekhari karthik sridharan wen sun runzhe wu bayesian active meta learning under prior misspecification link sabina j sloman ayush bharti samuel kaski contextual set selection under human feedback with model misspecification link shuo yang rajat sen sujay sanghavi building community driven libraries of natural programs link leonardo hernandez cano yewen pu robert d hawkins joshua b tenenbaum armando solar lezama modeled cognitive feedback to calibrate uncertainty for interactive learning link jaelle scheuerman zachary bishof chris j michael improving bionic limb control through batch reinforcement learning in an interactive game environment link kilian freitag rita laezza jan zbinden max ortiz catalan rewarded soups towards pareto optimality by interpolat...
|