Meta tags:
Headings (most frequently used words):
sunday, september, of, monday, 2016, 2010, 2009, blogs, by, ai, academics, rl, to, is, on, and, science, reinforcement, learning, june, 24, 2012, friday, 10, august, 16, july, 12, march, 22, contributors, resources, other, links, upcoming, conferences, blog, archive, why, deep, exciting, three, waves, great, talk, are, humans, just, another, primate, know, everything, visual, representation, doctorate, data, scientists, what, they, the, public, think, them, university, somehow, more, pure, than, industry,
Text of the page (most frequently used words):
the (67), and (47), that (25), this (24), #learning (22), share (21), from (15), for (12), with (10), ais (10), research (9), are (9), more (9), has (9), satinder (8), singh (8), deep (7), pinterest (7), facebook (7), blogthis (7), email (7), posted (7), not (7), what (7), into (7), humans (7), wave (7), comments (6), data (6), building (6), systems (6), category (6), these (6), new (6), will (6), september (5), exciting (5), often (5), can (5), they (5), level (5), but (5), about (5), skills (5), there (5), well (5), behavior (5), algorithms (5), 2016 (4), other (4), science (4), here (4), many (4), industry (4), sunday (4), etc (4), have (4), their (4), non (4), work (4), knowledge (4), such (4), academia (4), reinforcement (4), been (4), use (4), may (3), june (3), 2009 (3), 2010 (3), blog (3), academics (3), theory (3), posts (3), view (3), university (3), driven (3), researchers (3), associated (3), success (3), academic (3), generate (3), results (3), than (3), interesting (3), some (3), them (3), world (3), know (3), whether (3), human (3), build (3), machines (3), those (3), talk (3), logic (3), write (3), falls (3), decades (3), within (3), out (3), defined (3), need (3), renewed (3), focus (3), representations (3), function (3), memory (3), exploit (3), march (2), july (2), august (2), 2012 (2), three (2), waves (2), why (2), blogs (2), machine (2), post (2), curiosity (2), structure (2), broken (2), find (2), most (2), might (2), far (2), people (2), read (2), think (2), doctorate (2), representation (2), visual (2), monday (2), eric (2), circumstance (2), when (2), all (2), through (2), always (2), something (2), one (2), another (2), above (2), everything (2), every (2), intelligence (2), perhaps (2), high (2), problem (2), solving (2), deal (2), current (2), approach (2), chimp (2), side (2), difference (2), recently (2), case (2), great (2), useful (2), expert (2), fall (2), based (2), planning (2), reasoning (2), inference (2), main (2), experts (2), down (2), much (2), general (2), purpose (2), also (2), probabilities (2), lot (2), both (2), applications (2), least (2), ideas (2), towards (2), next (2), supervised (2), large (2), scale (2), major (2), largely (2), separates (2), handcrafted (2), statistical (2), tasks (2), yet (2), goals (2), third (2), continually (2), moment (2), contexts (2), rapid (2), themselves (2), cognitive (2), architectures (2), managing (2), course (2), boundaries (2), continual (2), addition (2), architecture (2), parallel (2), computations (2), hopefully (2), borrow (2), upon (2), witness (2), interest (2), really (2), control (2), internal (2), actor (2), critic (2), skip (2), october, november, 2008, archive, nips, links, upcoming, conferences, paul, krugman, uofmichigan, umass, rlai, uofalberta, resources, yael, niv, rich, sutton, peter, stone, michael, littman, andrew, contributors, subscribe, atom, home, older, washington, writer, former, stem, cell, researcher, harvard, argues, misplaced, incentive, agree, two, relevant, quotes, his, argument, real, constant, battle, recognition, rewards, space, speaking, engagements, funding, autonomy, consequently, while, described, reality, messier, curiously, tend, pursue, trendiest, technologies, explore, topics, happen, generous, levels, support, moreover, since, determined, almost, exclusively, number, prestige, publications, incentives, exceedingly, powerful, encourage, investigators, see, patterns, exist, disregard, contradictory, observations, important, overvalue, preliminary, unreliable, embrace, conclusions, deserve, viewed, greater, skepticism, article, somehow, pure, pew, center, press, must, lots, poll, commentary, scientists, public, comment, amusing, means, schmidt, ceo, google, describes, information, available, our, fingertips, presumably, networked, mobile, devices, being, literally, common, parlance, shifting, call, looking, internets, knowing, innocuous, meaning, word, phrase, changing, huge, shift, culture, society, you, want, hear, say, statement, around, minutes, video, link, friday, discussions, colleagues, best, linguistic, abilities, builds, navigate, physically, manipulate, basic, social, simply, stated, state, effective, would, dog, point, incremental, qualitative, differences, between, whilst, mostly, degree, came, across, robert, sapolsky, recommend, listening, including, questions, end, whet, your, appetite, until, was, thought, mind, distinguishes, apparently, wonderfully, delivered, informative, just, primate, recent, meeting, heard, darpa, synthesis, worldview, speak, somewhat, surprise, found, reasonably, consistent, own, way, partition, version, probability, inclusion, criterion, rules, determine, skill, suitably, asbtracted, simulator, which, derive, example, dynamic, programming, mdps, using, manually, crafted, graphical, models, over, emphasis, shifted, mixing, ongoing, effort, industrial, arguably, crested, invested, creating, pushing, past, decade, excitement, because, increasing, availability, labeled, computation, resurgent, force, spreading, industries, had, confined, last, few, years, thanks, due, deepmind, alphago, significant, component, nature, involve, classification, regression, reward, maximization, reached, crest, take, shape, though, its, genesis, reaching, back, earliest, perform, contextual, adaptation, adapting, previous, context, outcomes, even, shot, hallmark, contrast, second, part, motivated, intrinsic, experience, broadly, unsupervised, play, role, finally, elaborate, currently, used, flexibly, growing, set, now, considerable, crosses, laid, indeed, sharply, nonetheless, believe, helps, full, disclosure, founder, company, name, quite, settled, among, folks, cogitai, inc, john, launchbury, friend, asks, different, unfortunately, come, called, obviously, points, made, ijcai, workshop, usual, approximation, subtle, consequential, asynchronous, distributed, clever, gpus, café, theano, torch, tensor, flow, get, stuff, added, incorporate, computer, volumes, dimensional, particular, freed, shackles, illustrative, empirical, successes, increasingly, spoken, textual, crucial, renewing, expanding, progress, perception, forms, attention, application, old, goal, intelligently, foundational, processing, turn, makes, easier, enter, field, don, necessarily, catch, couple, sophisticated, turns, were, designed, special, classes, domains, increased, policy, gradient, algorithmic, innovation, continues, recurrence, episodic, gating, standard, candidates, exploration, remarkable, increase, stability, resulting, approximators, overcome, fear, earlier, negative, collective, experiences, worst, theoretical, approximaton, collectively, intuitions, soon, visualize, look, learned, inadequate, redesign, elements, neural, network, generally, artificial, occasional, professorial, life, sidebar,
Text of the page (random words):
reinforcement learning skip to main skip to sidebar reinforcement learning a blog about reinforcement learning machine learning and more generally artificial intelligence with occasional posts about professorial life monday september 5 2016 why is deep rl exciting every so often a friend asks what is so exciting or different about what has unfortunately come to be called deep rl other than the obviously exciting new applications here are some points i made in my talk at the ijcai 2016 deep rl workshop 1 renewed focus on learning representations in addition to the usual focus on function approximation this is a subtle but consequential difference more of us visualize look at the representations learned and if we find them inadequate we redesign elements of the neural network collectively we are building intuitions and hopefully soon theory about architecture and representations remarkable increase in stability of the resulting function approximators overcome some of the fear from earlier negative collective experiences and worst case theoretical results on function approximaton and rl 2 renewed focus on cognitive architecture in addition to the standard rl candidates e g actor critic there is exploration of many new architectures the use of recurrence episodic memory gating etc 3 renewed interest in the most foundational of rl algorithms in a parallel to supervised learning turns out we may not need algorithms that were designed to exploit special structure in classes of domains witness the increased use of q learning td actor critic policy gradient etc of course algorithmic innovation continues this in turn makes it far easier for non rl people to enter the field and do interesting and useful work they don t necessarily have to catch up on a couple of decades of the more sophisticated rl algorithms 4 rl for control of internal processing control of memory read write as well other forms of attention are a really exciting application of rl towards the old ai goal of managing internal computations intelligently 5 we can borrow exploit and build upon really rapid progress in deep learning for perception this has been crucial for renewing and expanding interest in rl in particular it has freed rl from the shackles of illustrative empirical work witness successes in visual rl but also increasingly spoken textual rl 6 we can borrow incorporate exploit and build upon use of computer science to deal with large volumes of high dimensional data café theano torch tensor flow etc hopefully we will get a lot of rl stuff added to these asynchronous distributed parallel computations with clever use of memory and gpus to scale deep rl posted by satinder singh at 6 12 pm no comments email this blogthis share to x share to facebook share to pinterest sunday september 4 2016 three waves of ai at a recent meeting i heard a darpa synthesis from john launchbury a worldview so to speak of ai that somewhat to my surprise i found reasonably consistent with my own view of a useful way to partition ai work here is my version handcrafted ais expert systems fall into this category but so do logic based as well as probability based planning problem solving reasoning and inference systems the main inclusion criterion is that experts write down much of the knowledge and skills of such systems in expert systems rules determine the behavior skill in planning systems experts write down a suitably asbtracted world simulator from which general purpose algorithms derive behavior and so for example dynamic programming on mdps falls into this category reasoning and inference using general purpose algorithms on manually crafted graphical models also falls into this category over the decades emphasis has shifted from logic to probabilities and more recently to mixing logic and probabilities a lot of great ongoing ai effort in both industrial applications as well as in academic research falls into this category arguably at least in academia this wave has crested so that many of those invested in and creating these ideas are pushing them towards the next wave statistical data driven ais both supervised learning sl and reinforcement learning rl systems fall into this category in the past decade much of the excitement in ai within industry and within academia has been in sl because of the increasing availability of large scale labeled data and computation deep learning dl has been a major and resurgent force in spreading these ideas through many industries rl had largely been confined to academia but in the last few years has broken out into industry thanks largely due to the success of deepmind cf alphago what separates these ais from handcrafted ais is that there is a significant component of learning from data often this learning is statistical in nature what separates these ais from the next category is that they involve well defined tasks of classification regression or reward maximization we have not yet reached the crest of this wave within ai in industry and academia continual learning ais this new wave of ai is yet to take shape though it has its genesis reaching back decades in the earliest goals of ai these third wave ais will perform contextual adaptation and learning continually adapting from moment to moment what they know from previous contexts to generate behavior and continually learning new skills and knowledge from the current context and outcomes of their behavior rapid learning even one shot learning in new contexts will be a hallmark of such ais in contrast to the second wave there need not always be well defined tasks for such ais and so at least in part their behavior will be motivated by intrinsic goals associated with the learning of new knowledge and skills from the experience they themselves generate more broadly unsupervised learning ul will play a major role finally such ais will need more elaborate cognitive architectures than currently used in sl and rl for flexibly managing their growing set of knowledge and skills now of course there is considerable work that crosses the boundaries laid out above indeed the boundaries themselves are perhaps not all that sharply defined nonetheless i believe this view helps for full disclosure i am a co founder of a continual learning company cogitai inc p s the name for the third wave is not quite settled among ai folks posted by satinder singh at 6 43 pm no comments email this blogthis share to x share to facebook share to pinterest sunday june 24 2012 a great talk on are humans just another primate every so often i have discussions with my ai colleagues about whether the best research approach to building human level intelligence is to build machines that have the linguistic abilities of humans or perhaps the high level problem solving skills of humans or whether it is to builds machines that can navigate physically manipulate their world and deal with other machines at a basic social level more simply stated whether in the current state of ai the more effective research approach would be building a human or building a dog or a chimp those on the side of building a human often point to the non incremental qualitative differences between humans and non humans whilst those on the side of building a chimp think that there is mostly a difference of degree i came across this wonderfully delivered and informative talk by robert sapolsky that i recommend listening to including the questions at the end to whet your appetite until recently it was thought that a theory of mind distinguishes humans from non humans apparently it is not the case posted by satinder singh at 6 37 pm no comments email this blogthis share to x share to facebook share to pinterest friday september 10 2010 to know everything eric schmidt ceo of google describes the circumstance when all the information in the world is available at our fingertips presumably through always networked mobile devices etc as being a circumstance when we can literally know everything it is interesting that common parlance is shifting to call looking something up on the internets as knowing that something at one level this is innocuous in that the meaning of a word or phrase is changing but at another level this is a huge shift in culture and society if you want to hear eric say the statement above here is the link around 11 minutes into the video posted by satinder singh at 6 34 am no comments email this blogthis share to x share to facebook share to pinterest monday august 16 2010 a visual representation of a doctorate an amusing representation of what a doctorate means posted by satinder singh at 8 18 am 1 comment email this blogthis share to x share to facebook share to pinterest sunday july 12 2009 data on scientists and what they and the public think of them an interesting post from the pew research center for the people press a must read for academics and researchers lots of poll results and some commentary posted by satinder singh at 9 05 am 3 comments email this blogthis share to x share to facebook share to pinterest sunday march 22 2009 is university science somehow more pure than industry science in this washington post article the writer a former stem cell researcher at harvard argues that the view that university research science is curiosity driven is misplaced and that the incentive structure is broken i agree here are two relevant quotes from his argument that i find real university researchers are in a constant battle for recognition and the rewards associated with success research space speaking engagements funding and autonomy consequently while academic research is often described as curiosity driven the reality is messier as curiously many researchers tend to pursue the trendiest technologies and explore topics that happen to be associated with the most generous levels of research support moreover since academic success is determined almost exclusively by the number and prestige of research publications the incentives to generate results are exceedingly powerful and can encourage investigators to see patterns that may not exist to disregard contradictory observations that might be important to overvalue data that might be preliminary or unreliable and to embrace conclusions that deserve to be viewed with far greater skepticism posted by satinder singh at 4 52 pm no comments email this blogthis share to x share to facebook share to pinterest older posts home subscribe to posts atom contributors andrew ng michael littman peter stone rich sutton satinder singh yael niv blogs by ai academics machine learning theory rl resources rlai at uofalberta rl at umass rl at uofmichigan blogs by other academics paul krugman s blog links to upcoming conferences nips blog archive 2016 2 september 2 why is deep rl exciting three waves of ai 2012 1 june 1 2010 2 september 1 august 1 2009 2 july 1 march 1 2008 14 november 1 october 3 june 2 may 8
|