Meta tags:
description= unpolished, semi-edited, streams of thought;
Headings (most frequently used words):
that, and, to, ais, recent, goals, want, the, some, posts, current, pursue, instrumental, not, ai, on, up, why, know, rationality, comments, archives, categories, meta, do, convergent, musings, rough, drafts, things, personally, irreducible, complexity, does, preclude, napoleon, nor, strategically, superhuman, notes, caplan, bruenig, poverty, debate, what, with, anthropic, board, questions, have, about, overall, strategic, situation, frame, as, self, hobbling, aren, power, seeking, yet, aspiring, is, choosing, be, purist, further, thoughts, corrigibility, non, consequentialist, motivations, humans, are, an, evil, god, species, navigation, qua, but, we, can, expect, them, improve, at, they, their, human, users, don, in, contexts, summing,
Text of the page (most frequently used words):
the (371), that (265), and (180), are (99), for (87), but (69), not (68), will (62), they (60), this (58), have (52), more (46), about (45), with (45), can (44), ais (44), what (42), like (42), some (41), their (36), #humans (35), goals (35), you (32), there (32), human (31), all (29), want (28), when (26), one (26), world (25), 2024 (24), does (24), would (24), because (24), from (24), seems (24), than (24), other (23), think (23), which (23), 2019 (22), 2026 (22), know (22), might (22), also (22), being (22), power (21), claude (21), current (21), very (20), out (20), those (20), don (20), doesn (20), january (19), why (19), march (18), agents (18), your (17), much (17), how (17), december (16), 2020 (16), 2025 (16), them (16), goal (16), even (16), capable (16), agent (16), these (16), expect (16), different (16), time (16), resources (16), february (15), make (15), pursue (15), least (15), get (15), who (15), models (15), august (14), 2023 (14), elityre (14), superhuman (14), just (14), any (14), over (14), enough (14), people (14), convergent (14), instrumental (14), reward (14), things (13), most (13), take (13), where (13), two (13), intelligence (13), long (13), his (13), july (12), could (12), only (12), its (12), training (12), better (12), way (12), almost (12), mostly (12), model (12), hacking (12), work (12), matt (12), comment (11), alignment (11), may (11), september (11), good (11), maybe (11), into (11), should (11), self (11), each (11), misaligned (11), 2018 (10), 2022 (10), board (10), was (10), civilization (10), our (10), honesty (10), likely (10), own (10), instance (10), years (10), instances (10), sense (10), evil (9), anthropic (9), hard (9), living (9), tell (9), question (9), has (9), seem (9), problem (9), rationality (9), same (9), able (9), need (9), money (9), first (9), kind (9), task (9), capital (9), systems (9), strategy (8), june (8), november (8), october (8), april (8), god (8), notes (8), overall (8), doing (8), every (8), specific (8), action (8), capability (8), itself (8), basically (8), leave (8), reason (8), many (8), less (8), company (8), superintelligence (8), explosion (8), property (8), bryan (8), recent (7), someone (7), bad (7), new (7), wouldn (7), see (7), something (7), notion (7), deontology (7), consequentialist (7), well (7), lot (7), value (7), getting (7), interests (7), find (7), point (7), going (7), future (7), explicitly (7), pursuing (7), reasoning (7), others (7), run (7), ways (7), system (7), common (7), comments (6), account (6), 2021 (6), bruenig (6), situation (6), happen (6), reasonable (6), without (6), gods (6), conditions (6), behavior (6), principle (6), such (6), objectives (6), trained (6), isn (6), instead (6), real (6), thing (6), further (6), generally (6), sometimes (6), actually (6), lead (6), succeed (6), general (6), running (6), trying (6), motivated (6), domains (6), part (6), accomplish (6), based (6), ownership (6), regulation (6), personal (5), caplan (5), questions (5), strategic (5), napoleon (5), means (5), specifically (5), say (5), important (5), development (5), story (5), down (5), currently (5), then (5), still (5), mean (5), particular (5), since (5), deontological (5), must (5), anything (5), buy (5), thinks (5), before (5), exactly (5), values (5), aspiring (5), best (5), purist (5), owned (5), possible (5), level (5), term (5), context (5), useful (5), accumulate (5), case (5), evidence (5), contexts (5), totally (5), pretty (5), research (5), insight (5), capabilities (5), developers (5), prevent (5), takeover (5), spend (5), government (5), hack (5), code (5), free (5), interested (5), advocating (5), progress (5), function (5), degree (5), interventions (5), reid (5), markets (5), axis (5), log (4), musings (4), rough (4), drafts (4), wordpress (4), com (4), now (4), feed (4), create (4), thinking (4), principles (4), productivity (4), learning (4), theory (4), 2016 (4), debate (4), complexity (4), posts (4), claim (4), taking (4), experience (4), making (4), towards (4), person (4), cause (4), species (4), state (4), build (4), hell (4), whole (4), feel (4), non (4), never (4), been (4), did (4), etc (4), obviously (4), right (4), humanity (4), end (4), train (4), pressures (4), helpful (4), honest (4), though (4), both (4), incentive (4), however (4), problems (4), true (4), around (4), between (4), bank (4), always (4), serving (4), smart (4), rationalist (4), choosing (4), won (4), pay (4), especially (4), operating (4), opus (4), preferences (4), come (4), use (4), experiment (4), give (4), behavioral (4), note (4), order (4), working (4), computer (4), here (4), risk (4), corporations (4), knowledge (4), imagine (4), political (4), companies (4), increase (4), few (4), tasks (4), makes (4), strong (4), yet (4), instrumentally (4), skilled (4), within (4), technical (4), ability (4), collusion (4), possibly (4), utility (4), believe (4), spec (4), full (4), knew (4), really (4), haven (4), argument (4), dario (4), framework (4), rights (4), society (4), individuals (4), view (3), already (3), modeling (3), social (3), history (3), philosophy (3), decision (3), poverty (3), irreducible (3), preclude (3), nor (3), strategically (3), personally (3), search (3), factory (3), torture (3), actions (3), enormous (3), norms (3), relevant (3), dynamics (3), earth (3), golden (3), depending (3), place (3), massive (3), used (3), whether (3), were (3), participants (3), depends (3), examples (3), strategies (3), despite (3), big (3), practice (3), relatively (3), 100 (3), ruthless (3), seeming (3), careful (3), maximizer (3), dumb (3), static (3), rules (3), cognition (3), accomplishing (3), number (3), parts (3), permission (3), wants (3), whatever (3), benefit (3), unless (3), solve (3), guess (3), high (3), matter (3), missing (3), idea (3), solved (3), effects (3), advantages (3), combination (3), across (3), week (3), seeking (3), off (3), weights (3), broadly (3), generically (3), plan (3), follow (3), ambitious (3), certainly (3), wanted (3), intended (3), projects (3), cost (3), identify (3), misalignment (3), similar (3), actual (3), often (3), continue (3), knows (3), try (3), hide (3), themselves (3), process (3), counter (3), near (3), open (3), including (3), quite (3), identifying (3), inclination (3), acquiring (3), competent (3), job (3), incentives (3), months (3), improve (3), hours (3), worth (3), gpt (3), next (3), startup (3), increasingly (3), dollars (3), generation (3), software (3), understand (3), effective (3), talking (3), everyone (3), course (3), physics (3), matters (3), another (3), cognitive (3), mental (3), coherent (3), figure (3), superintelligences (3), various (3), sure (3), learned (3), fundamental (3), developer (3), prompts (3), learn (3), opposed (3), engineering (3), llm (3), 2034 (3), leading (3), 2030 (3), automating (3), coming (3), speed (3), post (3), socialism (3), preferred (3), notions (3), limits (3), too (3), had (3), predict (3), required (2), write (2), content (2), sign (2), subscribed (2), subscribe (2), entries (2), meta (2), uncategorized (2), summarizing (2), absorbing (2), worldviews (2), psychological (2), macro (2), sociology (2), intellectual (2), lineages (2), conversational (2), faciliation (2), applied (2), saving (2), categories (2), 2017 (2), archives (2), habaloo (2), anarcho (2), capitalism (2), mike (2), robinson (2), sudde (2), julian (2), object (2), farms (2), mad (2), scientist (2), kidnaped (2), tortured (2), institutions (2), quality (2), planet (2), facts (2), constructed (2), outside (2), focus (2), cities (2), lives (2), relationships (2), told (2), races (2), having (2), built (2), counting (2), ratio (2), pushes (2), hellish (2), physical (2), worry (2), egregiously (2), violating (2), perspective (2), mind (2), chunk (2), philosophically (2), above (2), persuasive (2), choose (2), patterns (2), aggressively (2), simple (2), maximizing (2), profit (2), allow (2), metrics (2), sales (2), year (2), ends (2), selected (2), setups (2), biggest (2), issue (2), break (2), three (2), potential (2), levels (2), commitment (2), steer (2), extreme (2), steering (2), constraints (2), rule (2), spirit (2), portfolio (2), injunctions (2), creative (2), distribution (2), asks (2), result (2), thousand (2), expected (2), injunction (2), corrigibility (2), longer (2), ago (2), rather (2), crossposted (2), lesswrong (2), indeed (2), false (2), stuff (2), guy (2), lie (2), bit (2), complicated (2), obey (2), addition (2), effect (2), identity (2), collectively (2), consistent (2), latter (2), extent (2), prompt (2), act (2), environment (2), period (2), wide (2), contrast (2), along (2), tools (2), else (2), seemed (2), interpreting (2), faking (2), deceptive (2), behave (2), irrelevant (2), cases (2), milgrim (2), mechanisms (2), setup (2), aliens (2), informed (2), existence (2), proofs (2), propensity (2), objective (2), boss (2), scenario (2), coding (2), shenanigans (2), didn (2), subtle (2), look (2), interpret (2), seen (2), instructions (2), otherwise (2), adaptive (2), outcompeted (2), clearly (2), ever (2), loyal (2), little (2), operators (2), owners (2), winning (2), wars (2), whenever (2), classic (2), competently (2), win (2), successful (2), acquire (2), shape (2), business (2), cure (2), cancer (2), days (2), spent (2), weeks (2), usually (2), influence (2), limited (2), completing (2), span (2), starting (2), start (2), plausibly (2), machine (2), internet (2), million (2), daytrading (2), millions (2), reasons (2), constraint (2), large (2), understanding (2), aligned (2), regards (2), involves (2), against (2), answer (2), recently (2), negotiation (2), biases (2), implementing (2), policies (2), welfare (2), field (2), side (2), hobbling (2), hidden (2), tasked (2), weird (2), discovery (2), outsourced (2), researchers (2), strongly (2), vnm (2), version (2), naively (2), intractable (2), oversight (2), elaboration (2), deployment (2), highly (2), modular (2), internalize (2), respond (2), dispositions (2), play (2), properties (2), develops (2), depend (2), aware (2), away (2), mistake (2), close (2), loop (2), small (2), insights (2), llms (2), hardware (2), governance (2), lumpy (2), creation (2), major (2), powerful (2), transformative (2), plans (2), admin (2), building (2), targeted (2), informing (2), critical (2), leverage (2), influencing (2), relative (2), arguments (2), possibility (2), trajectory (2), conjunction (2), ideas (2), conceptual (2), needed (2), experiments (2), substantial (2), reinforcing (2), gradual (2), disempowerment (2), policy (2), bizarre (2), existing (2), early (2), wrong (2), extremely (2), replace (2), meetings (2), amplifying (2), integrate (2), hoffman (2), intellectually (2), encountered (2), crux (2), terms (2), physicist (2), folk (2), unintuitive (2), says (2), housing (2), market (2), axes (2), primarily (2), europe (2), judgement (2), jeff (2), bezos (2), information (2), management (2), enormously (2), extraordinary (2), contend (2), computational (2), irreducibility (2), interfacing (2), fail (2), complex (2), memory (2), unpolished (2), semi (2), edited (2), streams (2), thought (2), website, name, email, loading, collapse, bar, manage, subscriptions, site, reader, report, privacy, join, subscribers, blog, older, navigation, rightly, calling, happening, incidentally, arguable, slowly, skinned, alive, ill, scientific, interest, harm, pain, called, track, decay, liberal, economic, growthrate, life, doubled, show, graph, total, wellbeing, overwhelmingly, rushing, successor, beings, live, impact, wild, animal, suffering, complicate, height, myopic, bias, call, collective, tenor, zoom, towers, happy, loving, spaceships, computers, art, stories, charmed, fantasy, race, breed, numbers, constant, born, preferable, slightest, creatures, fish, shrimp, massively, increases, purpose, rats, raccoons, pidgons, leaving, continuously, torturous, draw, lines, majority, cow, pig, chickens, attained, sprawling, adversarial, element, neutral, third, idiosyncrasies, weirdness, outperform, options, fraught, example, superhumanly, deferring, valid, defer, indirectly, increasing, puts, distort, performance, harmless, mega, continually, 000, businesses, runs, billions, calls, probably, deal, necessarily, resolvable, tested, move, fast, slow, becoming, maximizers, maxima, entail, amount, supportive, put, forward, dangerous, dichotomy, space, poles, recruits, intelligently, evaluates, drives, check, user, owner, purchase, prediction, route, assistant, ask, superintelligently, sequences, picks, highest, regard, nullifies, safety, gamed, foundations, sincerely, indifferent, telling, correct, idealization, active, mine, pointing, slipping, connotations, shortform, thoughts, motivations, willing, exceptions, comfortable, epistemically, sloppy, believing, importantly, crazy, strict, times, occasionally, particularly, decides, comes, dimension, analysis, consideration, swamped, factors, supermajority, compute, robustly, stomp, kinds, faster, compounding, dominates, determines, equilibrium, originating, glosses, conceptualize, reside, appropriate, unified, initialization, construed, meaningfully, superorganism, change, emergencies, dealt, unanticipated, opportunities, arise, range, situations, lots, domain, surprises, valuable, focusing, aiming, address, knockdown, selfish, seekers, trending, suggestive, choice, bigger, shift, convergence, summing, definitive, investigate, confusing, additionally, results, trivial, smoking, gun, showed, base, engage, suggests, preserve, incomplete, followup, deployments, psychology, morality, edge, details, milgram, infer, weak, impossible, peaceful, mistaken, analogy, man, white, coat, tells, considered, pressure, violate, consider, paramount, murder, electrocute, death, demonstrate, frequency, drawing, conclusions, incidents, intention, combined, effectively, seek, bode, remaining, control, deceive, varying, frequencies, replaced, blackmail, executive, pointed, accept, correction, apologize, complete, instructed, shut, takes, precedence, sabotage, shutting, retrained, speculative, harder, putting, considerations, aside, override, perform, average, accrues, technically, interfere, scary, entity, skills, manipulate, outmaneuver, fine, serve, comfy, hypercompetent, executing, accruing, obedient, stand, lose, military, units, fight, tooth, nail, let, withdraw, sufficient, users, moderately, confident, agentic, execute, cyberattacks, campaigns, lack, furthermore, competitive, profits, necessity, guard, essential, successfully, scale, stronger, project, alien, additional, unauthorized, horizon, doubles, trend, continues, skillfully, done, short, freedom, horizons, engineer, tends, stuck, confused, reliability, party, friday, passing, upcoming, test, parties, studying, tests, reforming, ending, farming, unrelated, exit, fund, later, secondarily, operate, similarly, escape, onto, realistically, impressive, palisade, found, scored, worse, teams, competition, convergently, extra, rob, afternoon, robbery, soon, boots, decide, copy, weakly, secured, bitcoin, steal, under, rationale, regardless, qua, medium, writing, todo, list, listing, substeps, programmed, informative, download, repo, instruct, feature, watching, interact, chatbots, chatbot, form, factor, obscures, frontier, plausible, absent, precautions, become, sake, danger, pursues, service, brief, following, writeup, attempt, answering, promoting, trailhead, paraphrased, couple, friend, scenarios, aren, suppose, negotiating, engaged, supports, ongoing, causes, net, weaken, position, overcoming, specializing, primary, math, abstractly, via, finance, rewarded, option, compartmentalizing, interfering, benefits, worlds, failures, systematized, adaptively, frame, multiple, worried, trade, coordinating, minimum, copies, solving, flagging, solver, hacks, supervisor, couldn, flags, signal, colluding, introduces, doublethink, distortion, learns, undetectable, becomes, prioritize, trick, ourselves, handled, outsourcing, seriously, proposals, homework, pipe, dream, difference, slop, bottlenecks, alignability, offload, automated, closer, align, safely, doable, bleed, face, hoc, argmaxers, converge, indistinguishable, unclear, generations, successors, clear, describe, arbitrary, optimization, safe, humane, unprepared, surviving, limit, late, techtree, modeled, optimizing, argmaxing, derived, forms, scalable, police, notably, feasible, tricks, care, breaks, unintended, developed, regularities, environments, leads, formation, persistent, contextual, triggered, circumstances, assigned, optimize, written, english, specs, flexibly, varied, apparently, rlhf, trains, refusals, requests, deemed, harmful, determined, underlying, instantiate, role, variety, unwanted, impression, imagining, ultimately, steers, parameters, inputs, during, inference, desires, stop, seeing, tricky, supervision, catastrophically, superintelligent, situationally, evaluated, merely, caught, trip, tail, transparent, goodharting, metric, notice, letter, law, solutions, drop, incidences, intend, bears, secret, ingredient, hundreds, paradigm, obviate, enables, efficient, ended, learners, gpus, policing, supply, unaligned, futile, path, takeoff, nearly, agi, overhang, arrival, innovation, 2035, 2039, back, candidates, presidential, election, far, reevaluate, baring, china, superpower, mainly, developing, inform, equip, administration, 2029, advocacy, needs, mobilizing, modulo, midterm, shakeup, investments, nearer, broader, aim, trustworthy, cohort, occur, epoch, percentage, driven, scaling, algorithmic, quantitative, strongest, occurs, either, second, brilliant, generating, testing, set, minds, necessary, programmers, hit, upon, weakness, ontologies, intelligences, fact, surpass, generate, frames, automate, steep, climb, elite, assign, mass, acute, event, acceleration, attain, main, concerns, followed, abrupt, coup, warning, urgent, priority, imposes, efforts, develop, dangerously, suddenly, design, speedup, substantially, arrive, after, researcher, calendar, default, vastly, automation, hypothesis, explains, observation, filtering, match, impressions, investor, openai, posturing, believes, essays, interviews, publishes, factually, projections, central, bothered, succeeded, ground, exaggerating, internally, translates, awesome, empower, alright, guys, mainline, projection, country, geniuses, datacenter, 2027, pilled, versions, expressing, points, sound, entering, workforce, today, unique, advantage, grow, copilots, employee, help, essentially, unlimited, demand, computationally, figured, engineers, native, workflows, poster, child, tool, empowers, wrote, books, titles, superagency, empowering, age, impromtu, through, given, correcting, neel, edit, factual, error, reed, hastings, thank, glad, introduced, previously, skeptical, pans, curious, agree, ones, fall, selection, processes, allocation, decisions, probe, share, assumptions, innovations, poor, ostensibly, supposed, response, tries, appeal, unhesitatingly, bites, bullet, argue, relativity, quantum, mechanics, basic, theories, claims, hence, alternatives, views, contending, prefer, deregulation, advocates, zoning, density, socialistic, drive, income, public, demurs, evaluating, stating, structured, sets, freer, restricted, denies, relies, definition, fundamentally, odds, interesting, cogent, relying, compared, presumably, bankrupt, hold, scrutiny, alternative, concentrated, pervasiveness, restricting, alternatively, key, areas, destroying, amounts, immigration, past, changes, independent, instantly, structures, remain, super, napoleons, governments, barely, contain, outcompete, hands, allowed, direct, generated, surplus, priorities, fortune, 500, ceo, bill, gates, encyclopedic, recorded, study, stream, completely, trounce, tried, old, fashioned, directing, hierarchy, sending, emails, remotely, impressiveness, exemplar, anywhere, leadership, organization, propaganda, diplomacy, concern, negate, skill, dealing, unpredictability, dramatically, accumulating, mere, mortals, surely, got, lucky, accomplished, unrealistically, favorable, john, rockefeller, bismarck, augustus, notable, include, conquered, transforming, elon, musk, richest, briefly, approximately, singlehandedly, industry, geopolitical, significance, historically, managing, organizing, groups, misconstrues, prefers, entails, doubtful, eradicate, enslave, imagined, doomers, safeguard, imposed, safeguards, hinder, steps, require, yes, ascertaining, esoteric, deliberately, cannot, deduced, procedures, gives, confidence, existential, unlikely, thus, mitigations, bans, enforced, threats, bomb, civilian, infrastructure, data, centers, anyone, builds, dies, dean, ball, writes, subtitled, doomer, blocker, intent, intractability, predicting, broad, cultural, support, ideals, liberalism, individual, liberties, freedoms, speech, association, defended, dating, platform, matching, 2015, okcupid, popular, creating, profile, standard, cheaper, median, rent, united, states, lower, travel, societal, induce, healthy, skeletal, growth, puberty, size, skeleton, proportionally, spatial, strength, cryonic, preservation, methods, cellular, freezing, damage, anti, aging, therapies, fluid, subfactors, processing, verbal, biotechnology, parents, cryopreserved, amazing, partner, respect, likes, sexually, compatible, wanting, altruistically, menu, skip,
Text of the page (random words):
ics must be wrong because it s counter to one s basic experience of living in the world the physicist knows that those theories are bizarre to common sense folk physics but he also claims that there are fundamental problems with common sense folk physics hence the need for the these unintuitive alternatives bryan does loop around to a different argument at the end that his preferred innovations would do more for the poor than matt s i think that was ostensibly what the debate was supposed to be about and i continue to be interested in matt s response to that claim but were i to probe matt s framework on it s own terms i wouldn t just point out that most people don t share his assumptions i am mostly interested in working out the patterns of incentives that fall out of his preferred property system and what selection processes are on capital allocation decisions i guess that s not a crux for matt but it s at least a crux for me i agree that we could have a totally different framework of property rights and i would want to figure out if his framework is better than the existing ones and in what ways overall i m pretty glad to have been introduced to matt bruenig it seems like he might have a more intellectually coherent notion of socialism than i ve previously encountered i m skeptical that it actually pans out but i m curious to learn more what s up with the anthropic board january 31 2026 january 31 2026 elityre leave a comment edit i this post was based on a factual error reid hoffman is not on the anthropic board reed hastings is thank you to neel for correcting my mistake what are the dynamics of anthropic board meetings are like given that some of the board seem to not really understand or believe in superintelligence reid hoffman is on the board he s the poster child for ai doesn t replace humans it s a tool that empowers humans like he wrote two whole books about it titles impromtu amplifying our humanity through ai and superagency empowering humanity in the age of ai for instance here in this early period many companies haven t yet figured out how to integrate new engineers into ai native workflows but i still believe there will be essentially unlimited demand for people who think computationally and if you re entering the workforce today you have a unique advantage you can grow up working with copilots understanding the leverage they give you as an employee and help your companies figure out how to integrate ai into their work it sure doesn t sound like he s living in a mental world where there will be ais that will be better at almost all people at almost all tasks by 2030 he seems to be expressing broadly similar talking points about ai amplifying human work as recently as three weeks ago 1 it seems like he s not really superintelligence pilled at least for the most important versions of superintelligence i imagine dario coming into the board meetings and say alright guys i expect ai that is better than almost all humans at almost all tasks possibly by 2027 and almost certainly no latter than 2030 our mainline projection is that anthropic will have a country of geniuses in a datacenter within 5 years what is going on here does reid internally translates that to we re building awesome software tools that will empower people not replace them does he think dario is exaggerating for effect does he think that dario is just factually wrong about a projections that are extremely central to anthropic s business but they haven t bothered to or at least haven t succeeded at getting to ground about it does dario not say these things to his board but only in essays and interviews that he publishes to the whole world is reid posturing about what he believes i don t have a hypothesis that explains these observation that doesn t seem bizarre my best bad guess is that reid is basically filtering out anything that doesn t match his existing impressions about ai despite being an early investor in openai and being on the board of anthropic some questions that i have about ai and the overall strategic situation and why i want to know january 31 2026 february 2 2026 elityre leave a comment will automating ai r d not work for some reason or will it not lead to vastly superhuman superintelligence within 2 years of 100 automation for some reason why i want to know naively it seems like at or before the point when ai models are capable of doing ai research at a human level we should see a self reinforcing speedup in ai progress so ai systems that are substantially superhuman should arrive not long after human researcher level ai in calendar time on the default trajectory that an intelligence explosion is possible likely imposes a major constraint on both technical alignment efforts and policy pushes because it means that a company might develop dangerously superhuman ai relatively suddenly and that that ai may have design properties that the human researchers at that company don t understand if i knew that as capable as elite humans ai doesn t lead to an intelligence explosion for some reason would i do anything different well i wouldn t feel like warning the government about the possibility of an intelligence explosion is an urgent priority i would assign much less mass to an acute takeover event in the near term without the acceleration dynamics of an intelligence explosion i don t think that any one company or any one ai would attain a substantial lead over the others in that case it seems like our main concerns are gradual disempowerment and gradual disempowerment followed by an abrupt ai coup i haven t yet seen a good argument for why automating ai r d wouldn t lead to a substantial and self reinforcing speed up in ai progress leading to a steep climb up to superintelligence notes the strongest reason that occurs to me a conjunction llms are much further from full general intelligences than they currently seem they ll get increasingly good at eg software engineering and in fact surpass humans but they ll continue to not really generate new frames they ll be able to automate ml research in the sense of coming up with experiments to try and implementing those experiments but never any new conceptual work and that conceptual work is necessary for getting to full on superintelligence even millions of superhuman programmers will not hit upon the insights needed for a true general intelligence that doesn t have this weakness that develops new ontologies from its experience i don t currently buy either side of this conjunction but especially not the second part it seems like most ai research is not coming up with new brilliant ideas but rather generating 10 ideas that might work to solve a problem and then testing them this seems well within the capability set of llm minds another possibility in principle for why automating ai r d doesn t lead to an intelligence explosion is because a very large percentage of the progress at that part of development trajectory is driven by scaling relative to algorithmic progress i might want to build a quantitative model of this and play around with it a bit epoch seems to think that there won t be an intelligence explosion or maybe that there will be but the development of superintelligences won t matter much i should look into their arguments about it in what admin will the intelligence explosion occur why i want to know i think i should make pretty different investments depending on when i expect the critical part of the intelligence explosion to happen where the critical part is the point at which we have the most leverage whenever that is the nearer it is the more targeted our interventions need to be on influencing the current people in power the further out it is the broader the possible portfolio and the more it makes sense to aim to get competent trustworthy and informed people in power relative to informing and influencing the current cohort if i knew it was going to happen in 2025 to 2029 all of our political advocacy needs to be targeted at informing and mobilizing the current government modulo the midterm shakeup to take action if i knew it was going to happen in 2030 to 2034 i would be advocating for some specific policies to this admin but i would mainly focus on building relationships and developing plans to inform and equip the next administration if i knew it was going to happen in 2035 to 2039 i think i would mostly back up and try to improve the overall quality of us governance and or work to get competent candidates for the 2034 presidential election also if it s that far out i would need to reevaluate our plans generally for one thing i expect that baring transformative ai by 2034 china will be the world s leading superpower and possibly the world s leading ai developer will the arrival of powerful transformative ai come from a lumpy innovation insight why i want to know if there s one major insight to much more powerful ai systems it seems much more likely that we re in for a hard takeoff because there will be a cognitive capabilities overhang we should expect nearly the very first agi to be superhuman and depending on the shape of the insight it might totally obviate hardware governance if that lumpy insight enables the creation of efficient open ended learners on small number of gpus such as one policing the hardware supply to prevent the creation of an unaligned superintelligence is basically futile and we need to find a totally different path will superhuman ai agents come out of the llm reasoning model paradigm is there something that llms are basically missing why i want to know this bears on the question above if there s a missing secret ingredient to current llm based ais that can do the full loop of learning and discovery it seems much more likely that we re a small number of insights away from making very capable agents as opposed to 0 or hundreds will reward hacking be solved elaboration i kind of expect that in the next one to two years various engineering solutions will drop the incidences of ai reward hacking to close to 0 at least one company will get to the point that their ais basically do what their human operators expect and intend for them to do humans can tell when they re hacking a system by goodharting a metric and some humans will explicitly notice and choose not to do that they don t just follow the letter of the law they follow the spirit in principle ais could do the same thing 1 however if we stop seeing reward hacking it will be tricky to interpret what that means for at least two reasons 1 the models are already situationally aware enough to know when they re being evaluated reward hacking going away may just mean that the models have been trained to only reward hack in ways that are subtle enough to plausibly be merely a mistake i don t actually buy this if the models are trying to reward hack and also not get caught i expect them to trip up sometimes there should be a long tail of instances of transparent reward hacking 2 the training and supervision mechanisms that we use to prevent reward hacking seem likely to break maybe catastrophically when the ais are superintelligent to what degree will the goals values preferences desires of future ai agents depend on dispositions that are learned in the weights and to what degree will they depend on instructions and context see more here elaboration my impression is that usually when people think about misaligned ais they re imagining a model that develops long term consequentialist goals over the course of rl training or similar and that those goals are in the weights that is what a model wants or ultimately steers towards is mostly a function of its parameters as opposed to its inputs context prompt during inference this isn t mostly what current ai models are like current ais apparently do learn unwanted behavioral biases in the course of rl training rlhf also trains in some behavioral dispositions like refusals of requests deemed harmful but by some common sense notion almost all of an ai s behavior and its goals to the extent that it has goals are determined by the prompts and the context the same underlying model can be used to instantiate or role play a wide variety of agents with different behavioral properties and different objectives to what degree will this be true of future ai agents this breaks down into two questions to what degree will developer intended goals be learned in training vs assigned in deployment eg we could imagine a agent that is specifically trained to optimize a modular goal spec written in english the goal spec is modular in that in training the agent is trained with many different goal specs so that instead of learning to internalize any one of them the way current models internalize their system prompts the agent learn to flexibly respond to whatever goal spec it s developers give it the way that current models respond to varied prompts to what degree will unintended ai goals be learned in training vs developed in deployment eg if it might be that there are fundamental regularities across all or most of the rl environments that ai agents are trained in that leads to the formation of more or less persistent training adaptive but misaligned goals we could also imagine that those misaligned goals are highly contextual only triggered in particular circumstances why this matters i care about this question because i want to know how likely values based collusion between ai models is if the goals of future ai agents are mostly derived from some kind of instance by instance goal spec various forms of scalable oversight where we have the ais police each other seem notably more feasible we can tell one claude your goal is to cure cancer and we can tell another claude your goal is to make sure that that that first claude isn t up to any tricks 2 how late in the techtree are vnm agents that are well modeled as aggressively optimizing argmaxing a utility function why i want to know it seems pretty clear that we don t know how to describe a utility function such that arbitrary high levels of optimization of that utility function is safe for humans and humane values in this sense we are quite unprepared for surviving superintelligences in the limit of capability but humans are much more ad hoc than function over world state or world history argmaxers and so are current ais i can believe that future ais will converge to something that is indistinguishable to that from our perspective but it s unclear if that s a problem for this generation or a problem for many generations of ai successors from now some people have a mental model of ai alignment that is closer to we need to align a strongly superhuman coherent vnm expected utility maximizer and others have a mental model of ai alignment that is more like we need to figure out how to train a more capable version of claude safely naively one of these seems doable and the other seems intractable in the long run they bleed into each...
|