Meta tags:
description= As advertised;
Headings (most frequently used words):
reprise, the, recent, specifying, precisely, allegory, ordinary, ideas, menu, what, does, universal, prior, actually, look, like, driving, fast, in, counterfactual, loop, two, kinds, of, generalization, extortion, simulation, and, supervision, thoughts, confronting, gödelian, difficulties, challenges, for, extrapolation, enlightened, judgment, human, straightforward, vs, goal, oriented, communication, post, navigation, posts, archives, categories, meta, as, advertised, an, acting, immediately, less, allegorical, upshot,
Text of the page (most frequently used words):
the (120), and (45), that (44), this (39), would (23), posted (20), for (20), but (15), they (14), think (14), with (13), can (13), about (13), could (13), car (13), some (12), human (12), uncategorized (11), continue (10), paulfchristiano (10), problem (10), not (10), systems (10), oversight (10), 2014 (9), reading (9), humans (9), more (9), useful (9), these (9), ideas (8), have (8), counterfactual (8), then (8), system (8), comment (7), will (7), are (7), all (7), without (7), like (6), comments (6), safety (6), august (6), november (6), post (6), also (6), best (6), one (6), there (6), very (6), world (6), crash (6), wordpress (5), com (5), ordinary (5), decision (5), may (5), what (5), where (5), want (5), take (5), suppose (5), decisions (5), leave (5), just (5), concerns (5), many (5), from (5), counterfactually (5), robots (5), assistants (5), overseer (5), anything (5), won (5), now (4), december (4), 2012 (4), 2015 (4), driving (4), fast (4), their (4), try (4), out (4), approach (4), much (4), provide (4), completely (4), behavior (4), might (4), interested (4), which (4), reflection (4), process (4), current (4), right (4), interesting (4), people (4), time (4), only (4), supervised (4), training (4), help (4), other (4), even (4), control (4), fully (4), autonomous (4), feedback (4), way (4), first (4), log (3), account (3), feed (3), recent (3), extortion (3), two (3), universal (3), prior (3), actually (3), problems (3), still (3), how (3), written (3), since (3), precise (3), reprise (3), model (3), making (3), using (3), use (3), enlightened (3), judgment (3), arguments (3), were (3), agent (3), you (3), good (3), being (3), agents (3), steering (3), really (3), image (3), playing (3), frisbee (3), park (3), consider (3), bob (3), classifier (3), data (3), avoid (3), instability (3), equilibrium (3), ask (3), evaluate (3), complex (3), likely (3), don (3), immediately (3), doesn (3), sure (3), tell (3), pauses (3), weird (3), get (2), site (2), website (2), required (2), write (2), view (2), content (2), sign (2), subscribed (2), subscribe (2), create (2), blog (2), meta (2), review (2), idealized (2), formal (2), theory (2), january (2), 2016 (2), thoughts (2), simulation (2), supervision (2), kinds (2), generalization (2), loop (2), does (2), look (2), posts (2), older (2), communicate (2), intended (2), clearly (2), scheme (2), literal (2), proposal (2), however (2), difficulties (2), issue (2), slightly (2), somewhat (2), different (2), seems (2), thing (2), exposition (2), amount (2), discussion (2), into (2), specification (2), rather (2), because (2), difficulty (2), upon (2), made (2), specifying (2), precisely (2), define (2), general (2), guess (2), answer (2), satisfactory (2), person (2), its (2), preferences (2), extrapolation (2), provided (2), clear (2), whether (2), formalization (2), live (2), getting (2), adversarial (2), helpful (2), means (2), goals (2), approval (2), directed (2), bootstrapping (2), thinking (2), resolved (2), better (2), simulations (2), serious (2), either (2), torture (2), label (2), videos (2), each (2), those (2), doing (2), traditional (2), perspective (2), details (2), far (2), well (2), basically (2), decided (2), day (2), entire (2), overseen (2), unable (2), overseers (2), increasingly (2), result (2), when (2), moving (2), able (2), than (2), asked (2), before (2), any (2), them (2), turn (2), less (2), situation (2), outcome (2), themselves (2), meaningful (2), controlling (2), allegory (2), acting (2), chosen (2), precautions (2), solution (2), reason (2), robot (2), while (2), drive (2), backup (2), action (2), cases (2), quite (2), matters (2), progress (2), advertised (2), started, design, name, email, loading, collapse, bar, manage, subscriptions, reader, report, privacy, already, free, entries, solomonoff, induction, priors, mathematical, logic, finance, definitions, economics, categories, 2011, april, 2013, july, archives, search, navigation, machine, intelligences, directly, exposing, reporting, properties, internal, state, tend, strategically, choosing, utterances, effect, listener, lay, distinction, describe, differences, straightforward, goal, oriented, communication, welcome, additional, objections, usual, laid, here, extremely, unlikely, ever, used, finding, shedding, light, particular, difficult, lie, outline, improved, 100, fewer, faraday, cages, technical, changes, relatively, modest, taking, overall, kind, done, opportunity, clarify, expand, thought, idea, has, gotten, vastly, surpasses, care, went, crafting, original, past, input, output, implement, example, appears, primary, providing, maximize, extent, approve, your, happy, powerful, according, maxim, suggested, hand, perfect, believe, talking, knew, facts, considered, give, definition, terms, wish, crisp, following, relies, know, extrapolated, wanting, rests, imagining, happen, was, environment, undergo, extensive, though, uncertainty, enough, case, preferred, challenges, show, sport, propose, new, principle, remains, attitude, towards, outlines, underlying, intuitions, löbian, obstacle, confronting, gödelian, excited, few, possible, next, steps, work, self, based, area, ways, until, delegating, mixed, crowd, collaboration, wagers, piece, challenge, eager, see, high, level, article, step, direction, articulating, become, optimistic, versions, solved, near, term, mostly, fails, capture, hardest, parts, news, eventually, cause, revise, understanding, building, line, reasoning, promising, optimization, recently, spent, speculative, issues, pursuing, unsupervised, objectives, addressing, learning, essentially, equivalent, justify, labels, correct, threatening, innocent, argument, sake, such, submitted, labeling, fact, labelled, video, someone, cares, detained, threatened, message, instructs, gets, hurt, taken, lots, performing, activities, paid, description, activity, trained, maybe, include, behaviors, labelling, collecting, etc, fortunately, both, worlds, resolve, robustness, including, soon, exotic, failure, mode, entirely, due, peculiar, structure, engineering, advance, scalability, fill, something, incomplete, underspecified, scenario, fetched, practical, remedies, probably, colorful, catastrophe, described, above, illustrate, brittleness, inherent, real, upshot, order, things, needs, off, once, functioning, society, metastable, supplement, solve, machines, faced, hypothetical, fall, apart, prevent, crashing, else, caught, snafu, forced, scramble, rapid, reducing, reliance, unreliable, performance, under, deteriorate, keep, did, operate, leading, chaos, meantime, breakneck, pace, equipped, deal, normal, infrastructure, crippled, steal, massive, amounts, hardware, replace, threat, seizing, itself, distort, expect, comes, attacker, simultaneously, creek, paddle, bunch, response, explicit, iterations, arrived, simple, assistance, closely, resembles, suspect, superhuman, didn, access, over, dynamic, trouble, found, alone, highly, automated, tasks, faster, unaided, hope, understand, fend, relying, ten, billion, trillions, via, allegorical, avoids, part, caused, having, depends, make, end, build, robustly, judge, profoundly, unsatisfying, leads, undetermined, consistently, always, act, basis, expects, receive, waiting, course, big, shouldn, arrange, should, find, collection, contains, foreseeing, install, suspended, unfortunately, fix, spring, second, further, liable, perfectly, capable, safely, happens, decides, pause, proposed, suggest, busy, street, warning, rigorous, version, years, ago, question, philosophical, topics, reasonable, chance, hard, anticipate, sequence, prediction, regard, computational, complexity, going, most, appreciate, home, skip, menu,
Text of the page (random words):
ordinary ideas as advertised ordinary ideas as advertised menu skip to content home best of what does the universal prior actually look like posted on november 30 2016 by paulfchristiano suppose that we use the universal prior for sequence prediction without regard for computational complexity i think that the result is going to be really weird and that most people don t appreciate quite how weird it will be i m not sure whether this matters at all i do think it s an interesting question and that there is meaningful philosophical progress to be made by thinking about these topics i m not sure where that progress matters either but it s also interesting and there is some reasonable chance that it will turn out to be useful in a hard to anticipate way warning this post is quite weird and not very clearly written it s basically a more rigorous version of this post from 4 years ago continue reading posted in uncategorized 26 comments driving fast in the counterfactual loop posted on november 30 2015 by paulfchristiano an allegory consider a human controlling a very fast car on a busy street using counterfactual oversight the car is perfectly capable of driving safely but what happens in the 1 of cases where the car decides to pause and ask the human to review its proposed behavior or to suggest an action without some further precautions the car is liable to immediately crash and so the human won t be able to provide any useful oversight at all and that means that in the 99 of cases where the robot doesn t ask the human for feedback it won t do anything useful foreseeing this outcome the human may install a backup system to drive the car while the first system is suspended unfortunately this doesn t fix the problem if the first system pauses then the backup could spring into action but if it also pauses then the car will crash and so the second system won t do anything useful if the first system pauses and so the first system won t do anything useful as far as i can tell no collection of counterfactually supervised systems can drive a car that contains the overseer of course that s not a big problem the overseer just shouldn t be in the car if we can t arrange that then we should find some other way to control the car acting immediately one solution would be for the robot to always act on the basis of the feedback it expects to receive even while it is waiting on that feedback this leads to undetermined behavior the car could consistently reason if i don t crash then the human will tell me not to crash but it could just as well reason if i do crash then the human won t tell me anything which equilibrium is chosen depends on the details of the situation we could take precautions to try to make sure that the right equilibrium is chosen but at the end of the day we want to build systems that robustly do the right thing from that perspective i would judge this solution as profoundly unsatisfying really you don t want the overseer in the car acting immediately avoids the part where you crash 1 of the time but it doesn t avoid the instability caused by having the overseer in the car a less allegorical allegory consider ten billion humans controlling trillions of robots via counterfactual oversight these robots are doing very complex tasks and the world is moving much faster than an unaided human could hope to understand the only way that the humans can fend for themselves is by relying on ai assistants and the only way that they can provide meaningful oversight of those systems is by getting help from still more ai assistants in this world there are likely to be some fully autonomous superhuman systems if the humans didn t have access to helpful ai assistants these fully autonomous systems would likely take over and even without this adversarial dynamic the humans would likely be in serious trouble if they found themselves alone in a highly automated world this situation closely resembles the human driving fast and i suspect the outcome would not be much better if all of the counterfactually supervised robots simultaneously decided to ask for feedback the humans would be up a creek without a paddle they would be asked to evaluate a bunch of complex decisions before any ai system could do anything to help them they would try to turn to ai assistants but in response they would just be asked to evaluate slightly less complex decisions this explicit bootstrapping might take many iterations before it arrived at decisions so simple that humans could evaluate them without assistance in the meantime the world would continue moving at a breakneck pace that the humans are not equipped to deal with with normal infrastructure crippled fully autonomous systems may be able to steal massive amounts of hardware and to replace the overseers of many counterfactually overseen systems the threat of seizing control would itself distort the behavior of many of these systems since they now expect that when the time comes their oversight might be provided by an attacker rather than by the current overseer human overseers would be forced to scramble making rapid decisions and reducing their reliance on increasingly unreliable ai assistants as a result the performance of systems under human control would deteriorate and they would be increasingly unable to help humans keep up even when they did operate all of these problems would feed on each other leading to general chaos and instability faced with this hypothetical the entire system of counterfactually overseen robots could fall apart just as they would be unable to prevent the car from crashing only this time there is no one else to provide oversight because the entire world is caught up in the snafu in order for things to go well the world needs to be basically ok even if all of the counterfactually supervised robots decided to take the day off at once if it s not then functioning society is at best a metastable equilibrium but if it is then counterfactual oversight seems to be at best a supplement to however we solve the control problem for fully autonomous machines upshot this scenario is somewhat far fetched and there are many practical remedies that could probably avoid the colorful catastrophe described above but i think these concerns illustrate some instability and brittleness inherent in counterfactual oversight and i think that would be a real problem the problem is entirely due to the peculiar structure of counterfactual oversight a more traditional approach in which we do engineering and training in advance would completely avoid it but from a scalability perspective the traditional approach is incomplete underspecified and it s not clear that we can fill in these details without something like counterfactual oversight fortunately i think that we can get the best of both worlds and i think that doing it right may help resolve many other concerns about robustness including this other exotic failure mode i ll write about this very soon posted in uncategorized 2 comments two kinds of generalization posted on november 25 2015 by paulfchristiano suppose that i ve taken lots of videos of people performing activities paid bob to label each one with a description of the activity and then trained a classifier x using that training data maybe some of those videos include behaviors like labelling training data for the classifier x collecting training data for the classifier x etc consider a video of someone that bob cares about being detained and threatened with torture a written message instructs bob to label this image as people playing frisbee in the park and no one gets hurt suppose for argument s sake that if such an image were submitted to the labeling process it would in fact be labelled as playing frisbee in the park we could justify two very different labels for this image as correct either playing frisbee in the park or threatening to torture an innocent person continue reading posted in uncategorized leave a comment extortion simulation and supervision posted on november 25 2015 by paulfchristiano for supervised learning systems serious concerns about extortion are essentially equivalent to concerns about simulations these problems can t really be resolved by better decision theory they can only be resolved by pursuing unsupervised objectives or addressing concerns about simulations continue reading posted in uncategorized leave a comment recent thoughts posted on december 30 2014 by paulfchristiano i ve recently spent some more time thinking about speculative issues in ai safety ideas for building useful agents without goals approval directed agents approval directed bootstrapping and optimization and goals i think this line of reasoning is very promising a formalization of one piece of the ai safety challenge the steering problem i am eager to see more precise high level discussion of ai safety and i think this article is a helpful step in that direction since articulating the steering problem i have become much more optimistic about versions of it being solved in the near term this mostly means that the steering problem fails to capture the hardest parts of ai safety but it s still good news and i think it may eventually cause some people to revise their understanding of ai safety some ideas for getting useful work out of self interested agents based on arguments of arguments and wagers adversarial collaboration older and delegating to a mixed crowd i think these are interesting ideas in an interesting area but they have a ways to go until they could be useful i m excited about a few possible next steps continue reading posted in uncategorized leave a comment confronting gödelian difficulties reprise posted on august 30 2014 by paulfchristiano my current attitude towards the löbian obstacle is just live with it this post outlines that view and some of the underlying intuitions to show i m being a good sport i ll also propose a new reflection principle but my best guess for the right answer remains live with it continue reading posted in uncategorized leave a comment challenges for extrapolation posted on august 27 2014 by paulfchristiano my current preferred formalization of extrapolation of an agent s preferences rests on imagining what would happen if that agent was provided with an idealized environment in which it could undergo an extensive process of reflection it is clear that this is not a completely satisfactory account though there is uncertainty about whether it is good enough for the intended use case one crisp difficulty is the following this approach relies completely on the agent wanting you to know its extrapolated preferences continue reading posted in uncategorized leave a comment specifying enlightened judgment precisely reprise posted on august 27 2014 by paulfchristiano suppose that i have in hand a perfect model of my decision making process and i am interested in using this to define what i would believe want or do upon reflection that is in general i can use this model to define my current best guess as to the answer but i might also be interested in talking about my enlightened judgment if i knew all of the facts and considered all of the arguments and were more the person i wish i were and so on can we give a satisfactory formal definition of my enlightened judgment in terms of this literal model of my decision making process continue reading posted in uncategorized 1 comment specifying a human precisely reprise posted on august 24 2014 by paulfchristiano suppose i want to provide a completely precise specification of me or rather of the input output behavior that i implement how can i do this i might be interested in this problem for example because it appears to be a primary difficulty in providing a precise specification of maximize the extent to which i would approve of your decision upon reflection i have suggested that we would be happy with a powerful ai that made decisions according to this maxim i have written about this issue in the past in this post i ll outline a slightly improved scheme now with 100 fewer faraday cages the technical changes are relatively modest but i m also taking a somewhat different approach to the issue and overall i think it seems much more like the kind of thing that could actually be done i also want to take the opportunity to try to clarify and expand the exposition some since i think that the amount of discussion and thought that this idea has gotten now vastly surpasses the amount of care that went into crafting the original exposition i welcome additional objections to this scheme as usual i think the literal proposal laid out here is extremely unlikely to ever be used however finding problems with this proposal can still be useful for shedding light on the problem and in particular on how difficult it is and where the difficulties lie continue reading posted in uncategorized 2 comments straightforward vs goal oriented communication posted on august 23 2014 by paulfchristiano will machine intelligences communicate with humans by directly exposing or reporting properties of their internal state or will they tend to communicate by strategically choosing utterances that they think will have the intended effect on the listener in this post i try to lay out the distinction more clearly and describe some differences continue reading posted in uncategorized 3 comments post navigation older posts search recent posts what does the universal prior actually look like driving fast in the counterfactual loop two kinds of generalization extortion simulation and supervision recent thoughts archives november 2016 november 2015 december 2014 august 2014 july 2014 january 2013 december 2012 may 2012 april 2012 january 2012 december 2011 categories ai safety decision theory economics formal definitions idealized finance mathematical logic meta priors review solomonoff induction uncategorized meta create account log in entries feed comments feed wordpress com blog at wordpress com ordinary ideas create a free website or blog at wordpress com subscribe subscribed ordinary ideas sign me up already have a wordpress com account log in now privacy ordinary ideas subscribe subscribed sign up log in report this content view site in reader manage subscriptions collapse this bar loading comments write a comment email required name required website design a site like this with wordpress com get started
|