Meta tags:
Headings (most frequently used words):
initialization, weight, contents, constant, random, miscellaneous, history, see, also, references, further, reading, lecun, glorot, he, orthogonal, fixup, others,
Text of the page (most frequently used words):
the (122), #initialization (75), learning (65), and (61), neural (47), displaystyle (40), networks (31), for (31), deep (30), network (28), with (26), weight (23), machine (23), layer (20), that (20), training (18), random (18), this (17), was (17), arxiv (17), conference (17), variance (16), edit (16), proceedings (15), activation (13), normalization (13), international (13), orthogonal (13), artificial (12), gradient (12), are (12), from (11), lecun (11), each (11), weights (11), bengio (10), residual (9), recurrent (9), yoshua (9), method (9), then (9), initializing (9), zero (9), using (8), systems (8), glorot (8), linear (8), designed (8), wikipedia (7), self (7), intelligence (7), bias (7), without (7), pmlr (7), 2017 (7), relu (7), used (7), parameters (7), such (7), matrix (7), times (7), initialized (7), distribution (7), article (7), text (6), may (6), page (6), has (6), isbn (6), generative (6), convolutional (6), models (6), model (6), backpropagation (6), initialize (6), its (6), pass (6), called (6), during (6), values (6), other (6), matrices (6), add (5), last (5), all (5), data (5), lstm (5), vision (5), architectures (5), james (5), image (5), supervised (5), regression (5), history (5), strategies (5), batch (5), pre (5), fixup (5), unitary (5), connections (5), 2015 (5), since (5), can (5), semi (5), biases (5), trainable (5), contents (4), search (4), statistics (4), policy (4), use (4), articles (4), reliable (4), references (4), optimization (4), mlp (4), transformer (4), computer (4), john (4), david (4), yann (4), hinton (4), based (4), reasoning (4), diffusion (4), recognition (4), human (4), context (4), descent (4), doi (4), 2016 (4), 978 (4), scale (4), gradients (4), information (4), processing (4), martens (4), 2013 (4), pdf (4), xavier (4), jmlr (4), 2010 (4), problem (4), free (4), need (4), mean (4), field (4), theory (4), classification (4), also (4), example (4), forward (4), proposed (4), function (4), sometimes (4), tanh (4), mathbb (4), these (4), entries (4), fan (4), initial (4), main (4), hide (4), move (4), sidebar (4), toggle (3), view (3), you (3), inc (3), 2026 (3), short (3), description (3), wikidata (3), impact (3), autoencoder (3), perceptron (3), long (3), ian (3), goodfellow (3), ilya (3), sutskever (3), andrew (3), geoffrey (3), knowledge (3), list (3), large (3), general (3), agent (3), automated (3), algorithm (3), approach (3), symbolic (3), reinforcement (3), engineering (3), datasets (3), convolution (3), quasi (3), newton (3), clustering (3), parameter (3), 2021 (3), courville (3), aaron (3), mit (3), press (3), samuel (3), 2018 (3), advances (3), momentum (3), workshop (3), unsupervised (3), help (3), balduzzi (3), what (3), walk (3), feedforward (3), xiao (3), 2020 (3), better (3), good (3), how (3), train (3), layers (3), kernel (3), kaiming (3), 1998 (3), way (3), become (3), dependent (3), tuning (3), methods (3), like (3), not (3), output (3), though (3), they (3), alpha (3), mathrm (3), cases (3), normal (3), when (3), convergence (3), initializes (3), shape (3), one (3), first (3), only (3), sqrt (3), standard (3), related (3), sampling (3), factor (3), samples (3), entry (3), mathcal (3), uniform (3), independently (3), preserve (3), gate (3), neurons (3), learn (3), vector (3), describes (3), tools (3), languages (2), table (2), safety (2), contact (2), about (2), privacy (2), terms (2), hidden (2), categories (2), cs1 (2), maint (2), periodical (2), lacking (2), retrieved (2), applications (2), visual (2), art (2), video (2), architecture (2), act (2), gan (2), mamba (2), rnn (2), cnn (2), multilayer (2), state (2), unit (2), gru (2), memory (2), turing (2), quoc (2), alex (2), fei (2), frank (2), rosenblatt (2), rule (2), semantic (2), ibm (2), language (2), speech (2), synthesis (2), alexnet (2), protocol (2), neuro (2), open (2), source (2), rlhf (2), sarsa (2), rectifier (2), sigmoid (2), tradeoff (2), functions (2), projects (2), glossary (2), review (2), springer (2), 1007 (2), adaptive (2), computation (2), cambridge (2), 262 (2), 03561 (2), further (2), reading (2), performance (2), 35th (2), curran (2), associates (2), 1806 (2), understanding (2), george (2), sparse (2), pascal (2), wise (2), thirteenth (2), foundations (2), normalizing (2), eds (2), connectionism (2), perspective (2), frean (2), marcus (2), leary (2), lennox (2), lewis (2), kurt (2), wan (2), duo (2), mcwilliams (2), brian (2), shattered (2), resnets (2), answer (2), question (2), very (2), link (2), cite (2), icml (2), hessian (2), through (2), zhang (2), 2019 (2), cvpr (2), init (2), 1511 (2), lechao (2), sohl (2), dickstein (2), jascha (2), schoenholz (2), pennington (2), jeffrey (2), cnns (2), saxe (2), orr (2), genevieve (2), müller (2), klaus (2), robert (2), 2024 (2), 540 (2), 49430 (2), dying (2), examples (2), computational (2), physics (2), empirical (2), jaitly (2), units (2), vanishing (2), see (2), there (2), between (2), careful (2), decrease (2), minibatch (2), less (2), important (2), backward (2), directly (2), work (2), still (2), before (2), trained (2), specifically (2), makes (2), point (2), lambda (2), approx (2), 0507 (2), 6733 (2), selu (2), end (2), 7159 (2), miscellaneous (2), allow (2), frac (2), performs (2), order (2), small (2), others (2), similarly (2), scalar (2), every (2), branch (2), much (2), problems (2), improve (2), kernels (2), done (2), uniformly (2), top (2), out (2), same (2), two (2), means (2), usually (2), some (2), nonzero (2), forget (2), avoiding (2), exploding (2), contains (2), setting (2), constant (2), both (2), step (2), curve (2), net (2), forest (2), anomaly (2), detection (2), bayes (2), structured (2), prediction (2), analysis (2), dimensionality (2), reduction (2), feature (2), shot (2), sources (2), appearance (2), upload (2), file (2), changes (2), links (2), read (2), log (2), create (2), account (2), donate (2), menu (2), topic, mobile, cookie, statement, developers, code, conduct, legal, contacts, disclaimers, available, under, additional, apply, site, agree, registered, trademark, non, profit, organization, wikimedia, foundation, creative, commons, attribution, sharealike, license, rendered, parsoid, edited, august, utc, empty, https, org, index, php, title, weight_initialization, oldid, 1368615473, category, workplace, warfare, military, games, marketing, chatbot, psychosis, healthcare, fiction, education, engine, explainable, environmental, competition, arms, race, anthropomorphism, winter, slop, literacy, infrastructure, effect, center, bubble, boom, social, economic, opposition, centers, propaganda, virtual, politician, regulation, precautionary, principle, nationalism, ethics, elections, takeover, alignment, government, cold, war, political, graph, gnn, adversarial, variational, vae, highway, echo, gated, term, vit, differentiable, françois, chollet, daniel, kokotajlo, jan, leike, mustafa, suleyman, schulman, aidan, gomez, noam, shazeer, ashish, vaswani, andrej, karpathy, silver, demis, hassabis, oriol, vinyals, krizhevsky, goodnight, graves, stephen, grossberg, lotfi, zadeh, jürgen, schmidhuber, hopfield, paul, werbos, seppo, linnainmaa, seymour, papert, joseph, weizenbaum, bernard, widrow, oliver, selfridge, herbert, simon, cliff, shaw, allen, newell, nathaniel, rochester, mccarthy, marvin, minsky, takeo, kanade, kunihiko, fukushima, shun, ichi, amari, claude, shannon, christopher, manning, von, neumann, walter, pitts, warren, sturgis, mcculloch, alan, people, yago, dbpedia, conceptnet, bases, opencog, lida, clarion, soar, cognitive, reasoners, procedural, logic, programs, inference, engines, expert, deductive, classifiers, robot, control, autogpt, action, selection, muzero, driving, car, openai, five, alphazero, alphago, decisional, watsonx, watson, project, debater, oasis, genie, world, udio, suno, riffusion, music, generation, veo, seedance, sora, kling, hailuo, runway, gen, dream, stable, recraft, midjourney, imagen, ideogram, gpt, flux, firefly, dall, aurora, alphafold, facial, whisper, elevenlabs, ocr, hwr, wavenet, audio, implementations, physical, agent2agent, hypothetical, superintelligence, asi, agi, weak, lethal, autonomous, weapons, laws, humanity, exam, companion, intelligent, nmt, game, playing, theorem, proving, actor, critic, situated, sovereign, blended, vibe, coding, word, embedding, hallucination, recursive, improvement, reflection, llm, post, uncanny, valley, rag, adversary, autoregression, latent, imitation, prompt, augmentation, regularization, gating, softmax, batchnorm, attention, conjugate, sgd, overfitting, double, loss, hyperparameter, representation, constraint, satisfaction, planning, concepts, lists, software, proprietary, institutions, companies, algorithms, timeline, narkhede, meenal, bartakke, prashant, sutaone, mukul, june, science, business, media, llc, 322, 0269, 2821, issn, s10462, 021, 10033, 291, mass, brock, soham, smith, simonyan, karen, high, 2102, 06171, balles, lukas, hennig, philipp, 413, 1705, 07774, 404, dissecting, adam, sign, magnitude, stochastic, bjorck, nils, gomes, carla, selman, bart, weinberger, kilian, 02375, dahl, 1147, 1139, 30th, importance, bordes, antoine, 2011, 323, 315, fourteenth, lamblin, popovici, dan, larochelle, hugo, 2006, greedy, erhan, dumitru, vincent, 208, 201, why, does, 2009, 127, 1561, 2200000006, trends, klambauer, günter, unterthiner, thomas, mayr, andreas, hochreiter, sepp, 1989, pfeifer, schreter, fogelman, steels, amsterdam, elsevier, university, zurich, october, 1988, generalization, design, 1702, 08591, sussillo, abbott, 2014, 1412, 6558, journal, madison, usa, omnipress, 742, 60558, 907, 735, 27th, via, huang, shi, perez, felipe, jimmy, volkovs, maksims, 4483, 4475, 37th, improving, hongyi, dauphin, tengyu, 1901, 09321, xie, xiong, jiang, shiliang, ieee, pattern, 6185, 6176, beyond, exploring, solution, extremely, orthonormality, modulation, mishkin, dmytro, matas, jiri, 06422, henaff, mikael, szlam, arthur, tasks, 1602, 06662, arjovsky, martin, shah, amar, 1128, 06464, 1120, 33rd, evolution, bahri, yasaman, 5402, 05393, 5393, dynamical, isometry, 000, vanilla, 350, 342, 34th, ballard, andy, desjardins, guillaume, swirszcz, grzegorz, dalibard, valentin, rapid, skip, shaping, 2110, 01765, wei, provable, benefit, optimizing, 2001, 05992, mcclelland, ganguli, surya, exact, solutions, nonlinear, dynamics, 1312, 6120, xiangyu, ren, shaoqing, sun, jian, delving, into, rectifiers, surpassing, level, imagenet, 1502, 01852, kumar, siddharth, krishna, 1704, 08863, 256, 249, difficulty, bottou, leon, berlin, heidelberg, 8_2, tricks, trade, efficient, backprop, shin, yeonjong, yanhui, karniadakis, numerical, 1706, 4208, cicp, 0165, 1903, 06733, 1671, communications, massachusetts, jozefowicz, rafal, zaremba, wojciech, 2350, 2342, 32nd, exploration, navdeep, simple, rectified, 1504, 00941, tension, having, tradeoffs, causes, undesirable, trait, while, developed, automatically, tune, optimizers, generated, considerable, excitement, phase, possible, however, paper, demonstrated, well, chosen, hyperparameters, sufficient, needing, either, combination, 2010s, era, common, difficult, starting, bottom, contrastive, divergence, belief, early, described, perceptrons, attracting, fixed, recommend, randomly, begin, leq, particular, scaling, maps, interval, itself, thus, ensuring, overall, gain, around, operating, conditions, maximum, which, improves, hyperbolic, tangent, behave, any, let, concatenation, looks, norm, unbiased, moves, instead, subset, larger, total, transformers, multiplier, element, inside, branches, follows, introduction, allowed, deeper, than, previous, gave, rise, their, own, connection, could, stabilize, unnecessary, normalizations, vgg, been, generalized, fully, connected, orthonormal, proceeding, runs, divides, deviation, approximately, sequential, lsuv, proposes, parameterize, result, throughout, remain, found, sequence, modelling, odd, widths, heights, central, fill, illustration, filling, stride, padding, delta, iid, calculate, transpose, depending, whether, tall, wide, left, right, according, multiplied, depends, time, until, independent, depth, haar, measure, poorly, equal, identically, compromise, goals, continuous, popularized, activations, most, situations, multiplicative, signal, positive, value, likely, 305, typically, bounded, range, unbounded, cause, suggested, parts, identity, similar, idea, assume, unless, otherwise, stated, simplest, form, but, leads, causing, features, symmetry, kinds, where, number, discuss, specific, discussed, later, sections, note, even, titled, choice, affects, speed, within, signals, quality, final, proper, necessary, issues, saturation, creating, modified, assigning, outline, research, ijcai, iclr, neurips, iccv, emnlp, ecml, pkdd, eccv, aaai, journals, conferences, topological, statistical, pac, occam, risk, minimization, machines, mathematical, roc, confusion, coefficient, determination, diagnostics, mechanistic, interpretability, loop, crowdsourcing, active, humans, play, multi, temporal, difference, ecram, electrochemical, ram, memtransistor, spiking, informed, radiance, deepdream, lenet, som, restricted, boltzmann, reservoir, computing, esn, isolation, local, outlier, ransac, markov, conditional, graphical, sdl, sne, pgd, pca, nmf, lda, ica, cca, exploratory, shift, optics, dbscan, expectation, maximization, fuzzy, hierarchical, cure, birch, support, svm, relevance, rvm, logistic, naive, boosting, bagging, ensembles, decision, trees, apprenticeship, multimodal, ontology, grammar, induction, rank, association, rules, automl, cleaning, density, estimation, modeling, quantum, neuromorphic, curriculum, online, few, meta, transfer, paradigms, mining, part, series, remove, message, please, looking, more, unreliable, citations, challenged, removed, listed, technique, encyclopedia, item, printable, version, download, print, export, switch, legacy, parser, get, shortened, url, permanent, here, actions, english, talk, subsection, personal, special, pages, recent, community, portal, contribute, current, events, navigation, jump, content,
Text of the page (random words):
rning learning to rank grammar induction ontology learning multimodal learning supervised learning classification regression apprenticeship learning decision trees ensembles bagging boosting random forest k nn linear regression naive bayes artificial neural networks logistic regression perceptron relevance vector machine rvm support vector machine svm clustering birch cure hierarchical k means fuzzy expectation maximization em dbscan optics mean shift dimensionality reduction factor analysis exploratory cca ica lda nmf pca pgd t sne sdl structured prediction graphical models bayes net conditional random field hidden markov anomaly detection ransac k nn local outlier factor isolation forest neural networks autoencoder deep learning feedforward neural network recurrent neural network lstm gru esn reservoir computing boltzmann machine restricted gan diffusion model som convolutional neural network u net lenet alexnet deepdream neural field neural radiance field physics informed neural networks transformer vision mamba spiking neural network memtransistor electrochemical ram ecram reinforcement learning q learning policy gradient sarsa temporal difference td multi agent self play learning with humans active learning crowdsourcing human in the loop mechanistic interpretability rlhf model diagnostics coefficient of determination confusion matrix learning curve roc curve mathematical foundations kernel machines bias variance tradeoff computational learning theory empirical risk minimization occam learning pac learning statistical learning vc theory topological deep learning journals and conferences aaai cvpr eccv ecml pkdd emnlp iccv neurips icml iclr ijcai ml jmlr related articles glossary of artificial intelligence list of datasets for machine learning research list of datasets in computer vision and image processing outline of machine learning v t e in deep learning weight initialization or parameter initialization describes the initial step in creating a neural network a neural network contains trainable parameters that are modified during training weight initialization is the pre training step of assigning initial values to these parameters the choice of weight initialization method affects the speed of convergence the scale of neural activation within the network the scale of gradient signals during backpropagation and the quality of the final model proper initialization is necessary for avoiding issues such as vanishing and exploding gradients and activation function saturation note that even though this article is titled weight initialization both weights and biases are used in a neural network as trainable parameters so this article describes how both of these are initialized similarly trainable parameters in convolutional neural networks cnns are called kernels and biases and this article also describes these constant initialization edit we discuss the main methods of initialization in the context of a multilayer perceptron mlp specific strategies for initializing other network architectures are discussed in later sections for an mlp there are only two kinds of trainable parameters called weights and biases each layer l displaystyle l contains a weight matrix w l r n l 1 n l displaystyle w l in mathbb r n_ l 1 times n_ l and a bias vector b l r n l displaystyle b l in mathbb r n_ l where n l displaystyle n_ l is the number of neurons in that layer a weight initialization method is an algorithm for setting the initial values for w l b l displaystyle w l b l for each layer l displaystyle l the simplest form is zero initialization w l 0 b l 0 displaystyle w l 0 b l 0 zero initialization is usually used for initializing biases but it is not used for initializing weights as it leads to symmetry in the network causing all neurons to learn the same features in this page we assume b 0 displaystyle b 0 unless otherwise stated recurrent neural networks typically use activation functions with bounded range such as sigmoid and tanh since unbounded activation may cause exploding values le jaitly hinton 2015 1 suggested initializing weights in the recurrent parts of the network to identity and zero bias similar to the idea of residual connections and lstm with no forget gate in most cases the biases are initialized to zero though some situations can use a nonzero initialization for example in multiplicative units such as the forget gate of lstm the bias can be initialized to 1 to allow good gradient signal through the gate 2 for neurons with relu activation one can initialize the bias to a small positive value like 0 1 so that the gradient is likely nonzero at initialization avoiding the dying relu problem 3 305 4 random initialization edit random initialization means sampling the weights from a normal distribution or a uniform distribution usually independently lecun initialization edit lecun initialization popularized in lecun et al 1998 5 is designed to preserve the variance of neural activations during the forward pass it samples each entry in w l displaystyle w l independently from a distribution with mean 0 and variance 1 n l 1 displaystyle 1 n_ l 1 for example if the distribution is a continuous uniform distribution then the distribution is u 3 n l 1 displaystyle mathcal u pm sqrt 3 n_ l 1 glorot initialization edit glorot initialization or xavier initialization was proposed by xavier glorot and yoshua bengio 6 it was designed as a compromise between two goals to preserve activation variance during the forward pass and to preserve gradient variance during the backward pass for uniform initialization it samples each entry in w l displaystyle w l independently and identically from u 6 n l 1 n l 1 displaystyle mathcal u pm sqrt 6 n_ l 1 n_ l 1 in the context n l 1 displaystyle n_ l 1 is also called the fan in and n l 1 displaystyle n_ l 1 the fan out when the fan in and fan out are equal then glorot initialization is the same as lecun initialization he initialization edit as glorot initialization performs poorly for relu activation 7 he initialization or kaiming initialization was proposed by kaiming he et al 8 for networks with relu activation it samples each entry in w l displaystyle w l from n 0 2 n l 1 displaystyle mathcal n 0 2 n_ l 1 orthogonal initialization edit saxe et al 2013 9 proposed orthogonal initialization initializing weight matrices as uniformly random according to the haar measure semi orthogonal matrices multiplied by a factor that depends on the activation function of the layer it was designed so that if one initializes a deep linear network this way then its training time until convergence is independent of depth 10 sampling a uniformly random semi orthogonal matrix can be done by initializing x displaystyle x by iid sampling its entries from a standard normal distribution then calculate x x 1 2 x displaystyle left xx top right 1 2 x or its transpose depending on whether x displaystyle x is tall or wide 11 for cnn kernels with odd widths and heights orthogonal initialization is done this way initialize the central point by a semi orthogonal matrix and fill the other entries with zero as an illustration a kernel k displaystyle k of shape 3 3 c c displaystyle 3 times 3 times c times c is initialized by filling k 2 2 displaystyle k 2 2 with the entries of a random semi orthogonal matrix of shape c c displaystyle c times c and the other entries with zero balduzzi et al 2017 12 used it with stride 1 and zero padding this is sometimes called the orthogonal delta initialization 11 13 related to this approach unitary initialization proposes to parameterize the weight matrices to be unitary matrices with the result that at initialization they are random unitary matrices and throughout training they remain unitary this is found to improve long sequence modelling in lstm 14 15 orthogonal initialization has been generalized to layer sequential unit variance lsuv initialization it is a data dependent initialization method and can be used in convolutional neural networks it first initializes weights of each convolution or fully connected layer with orthonormal matrices then proceeding from the first to the last layer it runs a forward pass on a random minibatch and divides the layer s weights by the standard deviation of its output so that its output has variance approximately 1 16 17 fixup initialization edit in 2015 the introduction of residual connections allowed very deep neural networks to be trained much deeper than the 20 layers of the previous state of the art such as the vgg 19 residual connections gave rise to their own weight initialization problems and strategies these are sometimes called normalization free methods since using residual connection could stabilize the training of a deep neural network so much that normalizations become unnecessary fixup initialization is designed specifically for networks with residual connections and without batch normalization as follows 18 initialize the classification layer and the last layer of each residual branch to 0 initialize every other layer using a standard method such as he initialization and scale only the weight layers inside residual branches by l 1 2 m 2 displaystyle l frac 1 2m 2 add a scalar multiplier initialized at 1 in every branch and a scalar bias initialized at 0 before each convolution linear and element wise activation layer similarly t fixup initialization is designed for transformers without layer normalization 19 9 others edit instead of initializing all weights with random values on the order of o 1 n displaystyle o 1 sqrt n sparse initialization initialized only a small subset of the weights with larger random values and the other weights zero so that the total variance is still on the order of o 1 displaystyle o 1 20 random walk initialization was designed for mlp so that during backpropagation the l2 norm of gradient at each layer performs an unbiased random walk as one moves from the last layer to the first 21 looks linear initialization was designed to allow the neural network to behave like a deep linear network at initialization since w r e l u x w r e l u x w x displaystyle w mathrm relu x w mathrm relu x wx it initializes a matrix w displaystyle w of shape r n 2 m displaystyle mathbb r frac n 2 times m by any method such as orthogonal initialization then let the r n m displaystyle mathbb r n times m weight matrix to be the concatenation of w w displaystyle w w 22 miscellaneous edit for hyperbolic tangent activation function a particular scaling is sometimes used 1 7159 tanh 2 x 3 displaystyle 1 7159 tanh 2x 3 this was sometimes called lecun s tanh it was designed so that it maps the interval 1 1 displaystyle 1 1 to itself thus ensuring that the overall gain is around 1 in normal operating conditions and that f x displaystyle f x is at maximum when x 1 1 displaystyle x 1 1 which improves convergence at the end of training 23 5 in self normalizing neural networks the selu activation function s e l u x λ x if x 0 α e x α if x 0 displaystyle mathrm selu x lambda begin cases x text if x 0 alpha e x alpha text if x leq 0 end cases with parameters λ 1 0507 α 1 6733 displaystyle lambda approx 1 0507 alpha approx 1 6733 makes it such that the mean and variance of the output of each layer has 0 1 displaystyle 0 1 as an attracting fixed point this makes initialization less important though they recommend initializing weights randomly with variance 1 n l 1 displaystyle 1 n_ l 1 24 history edit random weight initialization was used since frank rosenblatt s perceptrons an early work that described weight initialization specifically was lecun et al 1998 5 before the 2010s era of deep learning it was common to initialize models by generative pre training using an unsupervised learning algorithm that is not backpropagation as it was difficult to directly train deep neural networks by backpropagation 25 26 for example a deep belief network was trained by using contrastive divergence layer by layer starting from the bottom 27 martens 2010 20 proposed hessian free optimization a quasi newton method to directly train deep networks the work generated considerable excitement that initializing networks without pre training phase was possible 28 however a 2013 paper demonstrated that with well chosen hyperparameters momentum gradient descent with weight initialization was sufficient for training neural networks without needing either quasi newton method or generative pre training a combination that is still in use as of 2024 29 since then the impact of initialization on tuning the variance has become less important with methods developed to automatically tune variance like batch normalization tuning the variance of the forward pass 30 and momentum based optimizers tuning the variance of the backward pass 31 there is a tension between using careful weight initialization to decrease the need for normalization and using normalization to decrease the need for careful weight initialization with each approach having its tradeoffs for example batch normalization causes training examples in the minibatch to become dependent an undesirable trait while weight initialization is architecture dependent 32 see also edit backpropagation normalization machine learning gradient descent vanishing gradient problem references edit le quoc v jaitly navdeep hinton geoffrey e 2015 a simple way to initialize recurrent networks of rectified linear units arxiv 1504 00941 cs ne jozefowicz rafal zaremba wojciech sutskever ilya 2015 06 01 an empirical exploration of recurrent network architectures proceedings of the 32nd international conference on machine learning pmlr 2342 2350 goodfellow ian bengio yoshua courville aaron 2016 deep learning adaptive computation and machine learning cambridge massachusetts the mit press isbn 978 0 262 03561 3 lu lu shin yeonjong su yanhui karniadakis george em 2019 dying relu and initialization theory and numerical examples communications in computational physics 28 5 1671 1706 arxiv 1903 06733 doi 10 4208 cicp oa 2020 0165 1 2 3 lecun yann bottou leon orr genevieve b müller klaus robert 1998 efficient backprop in orr genevieve b müller klaus robert eds neural networks tricks of the trade berlin heidelberg springer pp 9 50 doi 10 1007 3 540 49430 8_2 isbn 978 3 540 49430 0 retrieved 2024 10 05 glorot xavier bengio yoshua 2010 03 31 understanding the difficulty of training deep feedforward neural networks proceedings of the thirteenth international conference on artificial intelligence and statistics jmlr workshop and conference proceedings 249 256 kumar siddharth krishna 2017 on weight initialization in deep neural networks arxiv 1704 08863 cs lg he kaiming zhang xiangyu ren shaoqing sun jian 2015 delving deep into rectifiers surpassing human level performance on imagenet classification arxiv 1502 01852 cs cv saxe andrew m mcclelland james l ganguli surya 2013 exact solutions to the nonlinear dynamics of learning in deep linear neural networks ...
|