Meta tags:
Headings (most frequently used words):
actor, critic, algorithm, contents, overview, variants, see, also, references,
Text of the page (most frequently used words):
the (72), displaystyle (48), theta (37), learning (32), #critic (28), phi (25), gamma (25), actor (24), and (21), policy (18), gradient (16), function (16), reinforcement (13), algorithms (12), sum (12), value (11), advantage (10), edit (10), that (10), textstyle (10), this (9), neural (9), network (9), algorithm (8), with (8), then (8), leq (8), wikipedia (7), artificial (7), john (7), systems (7), action (7), text (6), page (6), machine (6), deep (6), method (6), can (6), lambda (6), nabla (6), are (6), left (6), right (6), state (5), based (5), control (5), bias (5), variance (5), for (5), arxiv (5), continuous (5), delta (5), some (5), contents (4), search (4), citations (4), from (4), intelligence (4), applications (4), reasoning (4), models (4), general (4), sarsa (4), estimation (4), methods (4), error (4), alpha (4), where (4), main (4), article (4), hide (4), move (4), sidebar (4), toggle (3), view (3), terms (3), april (3), 2026 (3), articles (3), short (3), wikidata (3), generative (3), optimization (3), david (3), silver (3), alex (3), knowledge (3), self (3), diffusion (3), image (3), synthesis (3), model (3), automated (3), source (3), descent (3), functions (3), hyperparameter (3), history (3), issn (3), doi (3), 978 (3), isbn (3), 2019 (3), high (3), generalized (3), asynchronous (3), references (3), also (3), step (3), learned (3), estimate (3), more (3), here (3), which (3), parameters (3), such (3), learn (3), cdot (3), reinforce (3), overview (3), actions (3), tools (3), languages (2), table (2), safety (2), contact (2), about (2), privacy (2), apply (2), using (2), use (2), was (2), last (2), categories (2), all (2), lacking (2), description (2), impact (2), visual (2), video (2), data (2), act (2), autoencoder (2), rnn (2), recurrent (2), vision (2), transformer (2), computer (2), turing (2), architectures (2), schulman (2), fei (2), andrew (2), graves (2), ibm (2), list (2), large (2), language (2), recognition (2), speech (2), human (2), implementations (2), protocol (2), agent (2), context (2), symbolic (2), open (2), improvement (2), training (2), projects (2), november (2), 2012 (2), 1109 (2), bibcode (2), ieee (2), survey (2), massachusetts (2), optimal (2), 2018 (2), mit (2), press (2), konda (2), vijay (2), tsitsiklis (2), lillicrap (2), timothy (2), soft (2), information (2), processing (2), 2017 (2), see (2), spaces (2), version (2), a2c (2), variants (2), returns (2), off (2), uses (2), exponentially (2), decaying (2), gae (2), estimating (2), parameterized (2), updated (2), both (2), let (2), common (2), higher (2), but (2), leftarrow (2), example (2), taken (2), respect (2), since (2), any (2), unbiased (2), estimators (2), these (2), baseline (2), psi (2), mathbb (2), goal (2), improve (2), reward (2), space (2), discrete (2), either (2), introducing (2), according (2), help (2), appearance (2), upload (2), file (2), changes (2), links (2), read (2), log (2), create (2), account (2), donate (2), menu (2), add, topic, mobile, cookie, statement, statistics, developers, code, conduct, legal, contacts, disclaimers, available, under, additional, may, site, you, agree, registered, trademark, non, profit, organization, wikimedia, foundation, inc, creative, commons, attribution, sharealike, license, rendered, parsoid, edited, utc, hidden, different, retrieved, https, org, index, php, title, critic_algorithm, oldid, 1348373706, category, workplace, warfare, military, art, games, marketing, chatbot, psychosis, healthcare, fiction, education, architecture, engine, explainable, environmental, competition, arms, race, anthropomorphism, winter, slop, literacy, infrastructure, effect, center, bubble, boom, social, economic, opposition, centers, propaganda, virtual, politician, regulation, precautionary, principle, nationalism, ethics, elections, takeover, alignment, government, cold, war, political, graph, gnn, adversarial, gan, variational, vae, mamba, highway, residual, convolutional, cnn, multilayer, perceptron, mlp, echo, gated, unit, gru, long, term, memory, lstm, vit, differentiable, françois, chollet, daniel, kokotajlo, jan, leike, mustafa, suleyman, aidan, gomez, noam, shazeer, ashish, vaswani, andrej, karpathy, demis, hassabis, ian, goodfellow, quoc, oriol, vinyals, ilya, sutskever, krizhevsky, james, goodnight, stephen, grossberg, lotfi, zadeh, yoshua, bengio, yann, lecun, jürgen, schmidhuber, hopfield, geoffrey, hinton, paul, werbos, seppo, linnainmaa, seymour, papert, joseph, weizenbaum, bernard, widrow, frank, rosenblatt, oliver, selfridge, herbert, simon, cliff, shaw, allen, newell, nathaniel, rochester, mccarthy, marvin, minsky, takeo, kanade, kunihiko, fukushima, shun, ichi, amari, claude, shannon, christopher, manning, von, neumann, walter, pitts, warren, sturgis, mcculloch, alan, people, yago, dbpedia, conceptnet, bases, opencog, lida, clarion, soar, cognitive, rule, semantic, reasoners, procedural, logic, programs, inference, engines, expert, deductive, classifiers, robot, autogpt, selection, muzero, driving, car, openai, five, alphazero, alphago, decisional, watsonx, watson, project, debater, oasis, genie, world, udio, suno, riffusion, music, generation, veo, seedance, sora, kling, hailuo, runway, gen, dream, stable, recraft, midjourney, imagen, ideogram, gpt, flux, firefly, dall, aurora, alphafold, facial, whisper, elevenlabs, ocr, hwr, wavenet, alexnet, audio, physical, agent2agent, hypothetical, superintelligence, asi, agi, weak, lethal, autonomous, weapons, laws, humanity, exam, companion, intelligent, nmt, game, playing, theorem, proving, situated, approach, neuro, sovereign, blended, vibe, coding, word, embedding, hallucination, recursive, reflection, supervised, rlhf, llm, post, uncanny, valley, rag, adversary, autoregression, latent, imitation, prompt, engineering, augmentation, datasets, regularization, weight, initialization, gating, rectifier, sigmoid, softmax, activation, batchnorm, normalization, convolution, attention, backpropagation, conjugate, quasi, newton, sgd, clustering, overfitting, double, tradeoff, regression, loss, parameter, representation, constraint, satisfaction, planning, concepts, lists, software, proprietary, institutions, companies, glossary, timeline, grondman, ivo, busoniu, lucian, lopes, gabriel, babuska, robert, 1307, 1094, 6977, tsmcc, 2218595, 2012ithms, 1291g, 1291, transactions, man, cybernetics, part, reviews, standard, natural, gradients, grossi, csaba, 2010, lectures, cham, springer, international, publishing, 031, 00423, bertsekas, dimitri, belmont, athena, scientific, 886529, sutton, richard, barto, adaptive, computation, series, cambridge, 262, 03924, introduction, january, 2003, 1166, 0363, 0129, 1137, s0363012901385691, 1143, siam, journal, hunt, jonathan, pritzel, alexander, heess, nicolas, erez, tom, tassa, yuval, wierstra, daan, 1509, 02971, haarnoja, tuomas, zhou, aurick, hartikainen, kristian, tucker, george, sehoon, tan, jie, kumar, vikash, zhu, henry, gupta, abhishek, 1812, 05905, moritz, philipp, jordan, michael, abbeel, pieter, 1506, 02438, dimensional, levine, sergey, mnih, volodymyr, badia, adrià, puigdomènech, mirza, mehdi, harley, tim, kavukcuoglu, koray, 2016, 1602, 01783, 1999, advances, arulkumaran, kai, deisenroth, marc, peter, brundage, miles, bharath, anil, anthony, brief, 1053, 5888, msp, 2743240, 2017ispm, 26a, 1708, 05866, signal, magazine, specialized, deterministic, ddpg, incorporates, entropy, maximization, improved, exploration, sac, parallel, a3c, introduces, smoothly, interpolates, between, monte, carlo, low, adjusted, pick, trade, average, being, decay, strength, similarly, maintains, denoted, temporal, difference, calculated, trained, although, train, just, positive, integer, lower, price, approx, simplest, trains, minimize, squared, rate, note, only, constitutes, moving, target, not, requires, stopping, point, automatic, differentiation, approximation, approximator, given, above, certain, appear, approximated, depend, must, alongside, known, obtained, infty, frac, arbitrary, detailed, there, many, linear, following, big, optimize, ascent, find, maximizes, expected, episodic, time, horizon, infinite, discount, factor, int, takes, argument, environment, produces, probability, distribution, while, estimates, combination, thereof, understood, over, pure, like, via, consists, two, components, determines, take, evaluates, those, work, cases, family, combine, iteration, includes, how, when, remove, message, please, precise, lacks, sufficient, corresponding, inline, free, encyclopedia, item, other, printable, download, pdf, print, export, switch, legacy, parser, get, shortened, url, cite, permanent, link, related, what, english, talk, українська, français, català, azərbaycanca, subsection, top, personal, special, pages, recent, community, portal, contribute, random, current, events, navigation, jump, content,
Text of the page (random words):
log in contents move to sidebar hide top 1 overview toggle overview subsection 1 1 actor 1 2 critic 2 variants 3 see also 4 references toggle the table of contents actor critic algorithm 4 languages azərbaycanca català français українська edit links article talk english read edit view history tools tools move to sidebar hide actions read edit view history general what links here related changes upload file permanent link page information cite this page get shortened url switch to legacy parser print export download as pdf printable version in other projects wikidata item appearance move to sidebar hide from wikipedia the free encyclopedia reinforcement learning algorithms this article includes a list of general references but lacks sufficient corresponding inline citations please help improve this article by introducing more precise citations april 2026 learn how and when to remove this message the actor critic algorithm ac is a family of reinforcement learning rl algorithms that combine policy based rl algorithms such as policy gradient methods and value based rl algorithms such as value iteration q learning sarsa and td learning 1 an ac algorithm consists of two main components an actor that determines which actions to take according to a policy function and a critic that evaluates those actions according to a value function 2 some ac algorithms are on policy some are off policy some apply to either continuous or discrete action spaces some work in both cases overview edit the actor critic methods can be understood as an improvement over pure policy gradient methods like reinforce via introducing a baseline actor edit the actor uses a policy function π a s displaystyle pi a s while the critic estimates either the value function v s displaystyle v s the action value q function q s a displaystyle q s a the advantage function a s a displaystyle a s a or any combination thereof the actor is a parameterized function π θ displaystyle pi _ theta where θ displaystyle theta are the parameters of the actor the actor takes as argument the state of the environment s displaystyle s and produces a probability distribution π θ s displaystyle pi _ theta cdot s if the action space is discrete then a π θ a s 1 displaystyle sum _ a pi _ theta a s 1 if the action space is continuous then a π θ a s d a 1 displaystyle int _ a pi _ theta a s da 1 the goal of policy optimization is to improve the actor that is to find some θ displaystyle theta that maximizes the expected episodic reward j θ displaystyle j theta j θ e π θ t 0 t γ t r t displaystyle j theta mathbb e _ pi _ theta left sum _ t 0 t gamma t r_ t right where γ displaystyle gamma is the discount factor r t displaystyle r_ t is the reward at step t displaystyle t and t displaystyle t is the time horizon which can be infinite the goal of policy gradient method is to optimize j θ displaystyle j theta by gradient ascent on the policy gradient j θ displaystyle nabla j theta as detailed on the policy gradient method page there are many unbiased estimators of the policy gradient θ j θ e π θ 0 j t θ ln π θ a j s j ψ j s 0 s 0 displaystyle nabla _ theta j theta mathbb e _ pi _ theta left sum _ 0 leq j leq t nabla _ theta ln pi _ theta a_ j s_ j cdot psi _ j big s_ 0 s_ 0 right where ψ j textstyle psi _ j is a linear sum of the following 0 i t γ i r i textstyle sum _ 0 leq i leq t gamma i r_ i γ j j i t γ i j r i textstyle gamma j sum _ j leq i leq t gamma i j r_ i the reinforce algorithm γ j j i t γ i j r i b s j textstyle gamma j sum _ j leq i leq t gamma i j r_ i b s_ j the reinforce with baseline algorithm here b displaystyle b is an arbitrary function γ j r j γ v π θ s j 1 v π θ s j textstyle gamma j left r_ j gamma v pi _ theta s_ j 1 v pi _ theta s_ j right td 1 learning γ j q π θ s j a j textstyle gamma j q pi _ theta s_ j a_ j γ j a π θ s j a j textstyle gamma j a pi _ theta s_ j a_ j advantage actor critic a2c 3 γ j r j γ r j 1 γ 2 v π θ s j 2 v π θ s j textstyle gamma j left r_ j gamma r_ j 1 gamma 2 v pi _ theta s_ j 2 v pi _ theta s_ j right td 2 learning γ j k 0 n 1 γ k r j k γ n v π θ s j n v π θ s j textstyle gamma j left sum _ k 0 n 1 gamma k r_ j k gamma n v pi _ theta s_ j n v pi _ theta s_ j right td n learning γ j n 1 λ n 1 1 λ k 0 n 1 γ k r j k γ n v π θ s j n v π θ s j textstyle gamma j sum _ n 1 infty frac lambda n 1 1 lambda cdot left sum _ k 0 n 1 gamma k r_ j k gamma n v pi _ theta s_ j n v pi _ theta s_ j right td λ learning also known as gae generalized advantage estimate 4 this is obtained by an exponentially decaying sum of the td n learning terms critic edit in the unbiased estimators given above certain functions such as v π θ q π θ a π θ displaystyle v pi _ theta q pi _ theta a pi _ theta appear these are approximated by the critic since these functions all depend on the actor the critic must learn alongside the actor the critic is learned by value based rl algorithms for example if the critic is estimating the state value function v π θ s displaystyle v pi _ theta s then it can be learned by any value function approximation method let the critic be a function approximator v ϕ s displaystyle v_ phi s with parameters ϕ displaystyle phi the simplest example is td 1 learning which trains the critic to minimize the td 1 error δ i r i γ v ϕ s i 1 v ϕ s i displaystyle delta _ i r_ i gamma v_ phi s_ i 1 v_ phi s_ i the critic parameters are updated by gradient descent on the squared td error ϕ ϕ α ϕ δ i 2 ϕ α δ i ϕ v ϕ s i displaystyle phi leftarrow phi alpha nabla _ phi delta _ i 2 phi alpha delta _ i nabla _ phi v_ phi s_ i where α displaystyle alpha is the learning rate note that the gradient is taken with respect to the ϕ displaystyle phi in v ϕ s i displaystyle v_ phi s_ i only since the ϕ displaystyle phi in γ v ϕ s i 1 displaystyle gamma v_ phi s_ i 1 constitutes a moving target and the gradient is not taken with respect to that this is a common source of error in implementations that use automatic differentiation and requires stopping the gradient at that point similarly if the critic is estimating the action value function q π θ displaystyle q pi _ theta then it can be learned by q learning or sarsa in sarsa the critic maintains an estimate of the q function parameterized by ϕ displaystyle phi denoted as q ϕ s a displaystyle q_ phi s a the temporal difference error is then calculated as δ i r i γ q θ s i 1 a i 1 q θ s i a i displaystyle delta _ i r_ i gamma q_ theta s_ i 1 a_ i 1 q_ theta s_ i a_ i the critic is then updated by θ θ α δ i θ q θ s i a i displaystyle theta leftarrow theta alpha delta _ i nabla _ theta q_ theta s_ i a_ i the advantage critic can be trained by training both a q function q ϕ s a displaystyle q_ phi s a and a state value function v ϕ s displaystyle v_ phi s then let a ϕ s a q ϕ s a v ϕ s displaystyle a_ phi s a q_ phi s a v_ phi s although it is more common to train just a state value function v ϕ s displaystyle v_ phi s then estimate the advantage by 3 a ϕ s i a i j 0 n 1 γ j r i j γ n v ϕ s i n v ϕ s i displaystyle a_ phi s_ i a_ i approx sum _ j in 0 n 1 gamma j r_ i j gamma n v_ phi s_ i n v_ phi s_ i here n displaystyle n is a positive integer the higher n displaystyle n is the more lower is the bias in the advantage estimation but at the price of higher variance the generalized advantage estimation gae introduces a hyperparameter λ displaystyle lambda that smoothly interpolates between monte carlo returns λ 1 displaystyle lambda 1 high variance no bias and 1 step td learning λ 0 displaystyle lambda 0 low variance high bias this hyperparameter can be adjusted to pick the optimal bias variance trade off in advantage estimation it uses an exponentially decaying average of n step returns with λ displaystyle lambda being the decay strength 4 variants edit asynchronous advantage actor critic a3c parallel and asynchronous version of a2c 3 soft actor critic sac incorporates entropy maximization for improved exploration 5 deep deterministic policy gradient ddpg specialized for continuous action spaces 6 see also edit reinforcement learning policy gradient method deep reinforcement learning references edit arulkumaran kai deisenroth marc peter brundage miles bharath anil anthony november 2017 deep reinforcement learning a brief survey ieee signal processing magazine 34 6 26 38 arxiv 1708 05866 bibcode 2017ispm 34 26a doi 10 1109 msp 2017 2743240 issn 1053 5888 konda vijay tsitsiklis john 1999 actor critic algorithms advances in neural information processing systems 12 mit press 1 2 3 mnih volodymyr badia adrià puigdomènech mirza mehdi graves alex lillicrap timothy p harley tim silver david kavukcuoglu koray 2016 06 16 asynchronous methods for deep reinforcement learning arxiv 1602 01783 1 2 schulman john moritz philipp levine sergey jordan michael abbeel pieter 2018 10 20 high dimensional continuous control using generalized advantage estimation arxiv 1506 02438 haarnoja tuomas zhou aurick hartikainen kristian tucker george ha sehoon tan jie kumar vikash zhu henry gupta abhishek 2019 01 29 soft actor critic algorithms and applications arxiv 1812 05905 lillicrap timothy p hunt jonathan j pritzel alexander heess nicolas erez tom tassa yuval silver david wierstra daan 2019 07 05 continuous control with deep reinforcement learning arxiv 1509 02971 konda vijay r tsitsiklis john n january 2003 on actor critic algorithms siam journal on control and optimization 42 4 1143 1166 doi 10 1137 s0363012901385691 issn 0363 0129 sutton richard s barto andrew g 2018 reinforcement learning an introduction adaptive computation and machine learning series 2 ed cambridge massachusetts the mit press isbn 978 0 262 03924 6 bertsekas dimitri p 2019 reinforcement learning and optimal control 2 ed belmont massachusetts athena scientific isbn 978 1 886529 39 7 grossi csaba 2010 algorithms for reinforcement learning synthesis lectures on artificial intelligence and machine learning 1 ed cham springer international publishing isbn 978 3 031 00423 0 grondman ivo busoniu lucian lopes gabriel a d babuska robert november 2012 a survey of actor critic reinforcement learning standard and natural policy gradients ieee transactions on systems man and cybernetics part c applications and reviews 42 6 1291 1307 bibcode 2012ithms 42 1291g doi 10 1109 tsmcc 2012 2218595 issn 1094 6977 v t e artificial intelligence ai history timeline glossary lists algorithms companies institutions projects software open source proprietary concepts automated reasoning automated planning constraint satisfaction knowledge representation parameter hyperparameter loss functions regression bias variance tradeoff double descent overfitting clustering gradient descent sgd quasi newton method conjugate gradient method backpropagation attention convolution normalization batchnorm activation softmax sigmoid rectifier gating weight initialization regularization datasets augmentation prompt engineering reinforcement learning q learning sarsa imitation policy gradient diffusion latent diffusion model autoregression adversary rag uncanny valley llm post training rlhf self supervised learning reflection recursive self improvement hallucination word embedding vibe coding blended ai open source ai sovereign ai symbolic ai neuro symbolic ai situated approach actor critic algorithm applications automated theorem proving general game playing machine learning in context learning artificial neural network deep learning language model large nmt reasoning model context protocol intelligent agent ai agent artificial human companion humanity s last exam lethal autonomous weapons laws generative ai weak ai hypothetical artificial general intelligence agi artificial superintelligence asi agent2agent protocol physical ai implementations audio visual alexnet wavenet human image synthesis hwr ocr computer vision speech synthesis 15 ai elevenlabs speech recognition whisper facial recognition alphafold text to image models aurora dall e firefly flux gpt image ideogram imagen midjourney recraft stable diffusion text to video models dream machine runway gen hailuo ai kling sora seedance veo music generation riffusion suno udio world models genie oasis text list of large language models project debater ibm watson ibm watsonx decisional alphago alphazero openai five self driving car muzero action selection autogpt robot control reasoning systems deductive classifiers expert systems inference engines knowledge based systems logic programs procedural reasoning systems semantic reasoners rule based systems cognitive architectures act r soar clarion lida opencog knowledge bases conceptnet wikidata dbpedia yago people alan turing warren sturgis mcculloch walter pitts john von neumann christopher d manning claude shannon shun ichi amari kunihiko fukushima takeo kanade marvin minsky john mccarthy nathaniel rochester allen newell cliff shaw herbert a simon oliver selfridge frank rosenblatt bernard widrow joseph weizenbaum seymour papert seppo linnainmaa paul werbos geoffrey hinton john hopfield jürgen schmidhuber yann lecun yoshua bengio lotfi a zadeh stephen grossberg alex graves james goodnight andrew ng fei fei li alex krizhevsky ilya sutskever oriol vinyals quoc v le ian goodfellow demis hassabis david silver andrej karpathy ashish vaswani noam shazeer aidan gomez john schulman mustafa suleyman jan leike daniel kokotajlo françois chollet neural network architectures neural turing machine differentiable neural computer transformer vision transformer vit recurrent neural network rnn long short term memory lstm gated recurrent unit gru echo state network multilayer perceptron mlp convolutional neural network cnn residual neural network rnn highway network mamba autoencoder variational autoencoder vae generative adversarial network gan graph neural network gnn political ai cold war ai in government ai safety alignment ai takeover elections ethics of ai eu ai act nationalism precautionary principle regulation of ai us virtual politician propaganda opposition to ai data centers social and economic ai boom ai bubble ai data center ai effect ai infrastructure ai literacy ai slop ai winter anthropomorphism arms race competition environmental impact explainable ai generative engine optimization in architecture in education in fiction in healthcare chatbot psychosis in marketing in video games in visual art military applications ai warfare workplace impact category retrieved from https en wikipedia org w index php title actor critic_algorithm oldid 1348373706 categories reinforcement learning machine learning algorithms artificial intelligence hidden categories articles with short description short description is different from wikidata articles lacking in text citations from april 2026 all articles lacking in text citations this page was last edited on 12 april 2026 at 08 53 utc page was rendered with parsoid text is available under the creative commons attribution sharealike 4 0 license additional terms may apply by using this site...
|