Meta tags:
Headings (most frequently used words):
initialization, weight, contents, constant, random, miscellaneous, history, see, also, references, further, reading, lecun, glorot, he, orthogonal, fixup, others,
Text of the page (most frequently used words):
the (122), #initialization (75), learning (65), and (61), neural (47), displaystyle (40), networks (31), for (31), deep (30), network (28), with (26), weight (23), machine (23), layer (20), that (20), training (18), random (18), this (17), was (17), arxiv (17), conference (17), variance (16), edit (16), proceedings (15), activation (13), normalization (13), international (13), orthogonal (13), artificial (12), gradient (12), are (12), from (11), lecun (11), each (11), weights (11), bengio (10), residual (9), recurrent (9), yoshua (9), method (9), then (9), initializing (9), zero (9), using (8), systems (8), glorot (8), linear (8), designed (8), wikipedia (7), self (7), intelligence (7), bias (7), without (7), pmlr (7), 2017 (7), relu (7), used (7), parameters (7), such (7), matrix (7), times (7), initialized (7), distribution (7), article (7), text (6), may (6), page (6), has (6), isbn (6), generative (6), convolutional (6), models (6), model (6), backpropagation (6), initialize (6), its (6), pass (6), called (6), during (6), values (6), other (6), matrices (6), add (5), last (5), all (5), data (5), lstm (5), vision (5), architectures (5), james (5), image (5), supervised (5), regression (5), history (5), strategies (5), batch (5), pre (5), fixup (5), unitary (5), connections (5), 2015 (5), since (5), can (5), semi (5), biases (5), trainable (5), contents (4), search (4), statistics (4), policy (4), use (4), articles (4), reliable (4), references (4), optimization (4), mlp (4), transformer (4), computer (4), john (4), david (4), yann (4), hinton (4), based (4), reasoning (4), diffusion (4), recognition (4), human (4), context (4), descent (4), doi (4), 2016 (4), 978 (4), scale (4), gradients (4), information (4), processing (4), martens (4), 2013 (4), pdf (4), xavier (4), jmlr (4), 2010 (4), problem (4), free (4), need (4), mean (4), field (4), theory (4), classification (4), also (4), example (4), forward (4), proposed (4), function (4), sometimes (4), tanh (4), mathbb (4), these (4), entries (4), fan (4), initial (4), main (4), hide (4), move (4), sidebar (4), toggle (3), view (3), you (3), inc (3), 2026 (3), short (3), description (3), wikidata (3), impact (3), autoencoder (3), perceptron (3), long (3), ian (3), goodfellow (3), ilya (3), sutskever (3), andrew (3), geoffrey (3), knowledge (3), list (3), large (3), general (3), agent (3), automated (3), algorithm (3), approach (3), symbolic (3), reinforcement (3), engineering (3), datasets (3), convolution (3), quasi (3), newton (3), clustering (3), parameter (3), 2021 (3), courville (3), aaron (3), mit (3), press (3), samuel (3), 2018 (3), advances (3), momentum (3), workshop (3), unsupervised (3), help (3), balduzzi (3), what (3), walk (3), feedforward (3), xiao (3), 2020 (3), better (3), good (3), how (3), train (3), layers (3), kernel (3), kaiming (3), 1998 (3), way (3), become (3), dependent (3), tuning (3), methods (3), like (3), not (3), output (3), though (3), they (3), alpha (3), mathrm (3), cases (3), normal (3), when (3), convergence (3), initializes (3), shape (3), one (3), first (3), only (3), sqrt (3), standard (3), related (3), sampling (3), factor (3), samples (3), entry (3), mathcal (3), uniform (3), independently (3), preserve (3), gate (3), neurons (3), learn (3), vector (3), describes (3), tools (3), languages (2), table (2), safety (2), contact (2), about (2), privacy (2), terms (2), hidden (2), categories (2), cs1 (2), maint (2), periodical (2), lacking (2), retrieved (2), applications (2), visual (2), art (2), video (2), architecture (2), act (2), gan (2), mamba (2), rnn (2), cnn (2), multilayer (2), state (2), unit (2), gru (2), memory (2), turing (2), quoc (2), alex (2), fei (2), frank (2), rosenblatt (2), rule (2), semantic (2), ibm (2), language (2), speech (2), synthesis (2), alexnet (2), protocol (2), neuro (2), open (2), source (2), rlhf (2), sarsa (2), rectifier (2), sigmoid (2), tradeoff (2), functions (2), projects (2), glossary (2), review (2), springer (2), 1007 (2), adaptive (2), computation (2), cambridge (2), 262 (2), 03561 (2), further (2), reading (2), performance (2), 35th (2), curran (2), associates (2), 1806 (2), understanding (2), george (2), sparse (2), pascal (2), wise (2), thirteenth (2), foundations (2), normalizing (2), eds (2), connectionism (2), perspective (2), frean (2), marcus (2), leary (2), lennox (2), lewis (2), kurt (2), wan (2), duo (2), mcwilliams (2), brian (2), shattered (2), resnets (2), answer (2), question (2), very (2), link (2), cite (2), icml (2), hessian (2), through (2), zhang (2), 2019 (2), cvpr (2), init (2), 1511 (2), lechao (2), sohl (2), dickstein (2), jascha (2), schoenholz (2), pennington (2), jeffrey (2), cnns (2), saxe (2), orr (2), genevieve (2), müller (2), klaus (2), robert (2), 2024 (2), 540 (2), 49430 (2), dying (2), examples (2), computational (2), physics (2), empirical (2), jaitly (2), units (2), vanishing (2), see (2), there (2), between (2), careful (2), decrease (2), minibatch (2), less (2), important (2), backward (2), directly (2), work (2), still (2), before (2), trained (2), specifically (2), makes (2), point (2), lambda (2), approx (2), 0507 (2), 6733 (2), selu (2), end (2), 7159 (2), miscellaneous (2), allow (2), frac (2), performs (2), order (2), small (2), others (2), similarly (2), scalar (2), every (2), branch (2), much (2), problems (2), improve (2), kernels (2), done (2), uniformly (2), top (2), out (2), same (2), two (2), means (2), usually (2), some (2), nonzero (2), forget (2), avoiding (2), exploding (2), contains (2), setting (2), constant (2), both (2), step (2), curve (2), net (2), forest (2), anomaly (2), detection (2), bayes (2), structured (2), prediction (2), analysis (2), dimensionality (2), reduction (2), feature (2), shot (2), sources (2), appearance (2), upload (2), file (2), changes (2), links (2), read (2), log (2), create (2), account (2), donate (2), menu (2), topic, mobile, cookie, statement, developers, code, conduct, legal, contacts, disclaimers, available, under, additional, apply, site, agree, registered, trademark, non, profit, organization, wikimedia, foundation, creative, commons, attribution, sharealike, license, rendered, parsoid, edited, august, utc, empty, https, org, index, php, title, weight_initialization, oldid, 1368615473, category, workplace, warfare, military, games, marketing, chatbot, psychosis, healthcare, fiction, education, engine, explainable, environmental, competition, arms, race, anthropomorphism, winter, slop, literacy, infrastructure, effect, center, bubble, boom, social, economic, opposition, centers, propaganda, virtual, politician, regulation, precautionary, principle, nationalism, ethics, elections, takeover, alignment, government, cold, war, political, graph, gnn, adversarial, variational, vae, highway, echo, gated, term, vit, differentiable, françois, chollet, daniel, kokotajlo, jan, leike, mustafa, suleyman, schulman, aidan, gomez, noam, shazeer, ashish, vaswani, andrej, karpathy, silver, demis, hassabis, oriol, vinyals, krizhevsky, goodnight, graves, stephen, grossberg, lotfi, zadeh, jürgen, schmidhuber, hopfield, paul, werbos, seppo, linnainmaa, seymour, papert, joseph, weizenbaum, bernard, widrow, oliver, selfridge, herbert, simon, cliff, shaw, allen, newell, nathaniel, rochester, mccarthy, marvin, minsky, takeo, kanade, kunihiko, fukushima, shun, ichi, amari, claude, shannon, christopher, manning, von, neumann, walter, pitts, warren, sturgis, mcculloch, alan, people, yago, dbpedia, conceptnet, bases, opencog, lida, clarion, soar, cognitive, reasoners, procedural, logic, programs, inference, engines, expert, deductive, classifiers, robot, control, autogpt, action, selection, muzero, driving, car, openai, five, alphazero, alphago, decisional, watsonx, watson, project, debater, oasis, genie, world, udio, suno, riffusion, music, generation, veo, seedance, sora, kling, hailuo, runway, gen, dream, stable, recraft, midjourney, imagen, ideogram, gpt, flux, firefly, dall, aurora, alphafold, facial, whisper, elevenlabs, ocr, hwr, wavenet, audio, implementations, physical, agent2agent, hypothetical, superintelligence, asi, agi, weak, lethal, autonomous, weapons, laws, humanity, exam, companion, intelligent, nmt, game, playing, theorem, proving, actor, critic, situated, sovereign, blended, vibe, coding, word, embedding, hallucination, recursive, improvement, reflection, llm, post, uncanny, valley, rag, adversary, autoregression, latent, imitation, prompt, augmentation, regularization, gating, softmax, batchnorm, attention, conjugate, sgd, overfitting, double, loss, hyperparameter, representation, constraint, satisfaction, planning, concepts, lists, software, proprietary, institutions, companies, algorithms, timeline, narkhede, meenal, bartakke, prashant, sutaone, mukul, june, science, business, media, llc, 322, 0269, 2821, issn, s10462, 021, 10033, 291, mass, brock, soham, smith, simonyan, karen, high, 2102, 06171, balles, lukas, hennig, philipp, 413, 1705, 07774, 404, dissecting, adam, sign, magnitude, stochastic, bjorck, nils, gomes, carla, selman, bart, weinberger, kilian, 02375, dahl, 1147, 1139, 30th, importance, bordes, antoine, 2011, 323, 315, fourteenth, lamblin, popovici, dan, larochelle, hugo, 2006, greedy, erhan, dumitru, vincent, 208, 201, why, does, 2009, 127, 1561, 2200000006, trends, klambauer, günter, unterthiner, thomas, mayr, andreas, hochreiter, sepp, 1989, pfeifer, schreter, fogelman, steels, amsterdam, elsevier, university, zurich, october, 1988, generalization, design, 1702, 08591, sussillo, abbott, 2014, 1412, 6558, journal, madison, usa, omnipress, 742, 60558, 907, 735, 27th, via, huang, shi, perez, felipe, jimmy, volkovs, maksims, 4483, 4475, 37th, improving, hongyi, dauphin, tengyu, 1901, 09321, xie, xiong, jiang, shiliang, ieee, pattern, 6185, 6176, beyond, exploring, solution, extremely, orthonormality, modulation, mishkin, dmytro, matas, jiri, 06422, henaff, mikael, szlam, arthur, tasks, 1602, 06662, arjovsky, martin, shah, amar, 1128, 06464, 1120, 33rd, evolution, bahri, yasaman, 5402, 05393, 5393, dynamical, isometry, 000, vanilla, 350, 342, 34th, ballard, andy, desjardins, guillaume, swirszcz, grzegorz, dalibard, valentin, rapid, skip, shaping, 2110, 01765, wei, provable, benefit, optimizing, 2001, 05992, mcclelland, ganguli, surya, exact, solutions, nonlinear, dynamics, 1312, 6120, xiangyu, ren, shaoqing, sun, jian, delving, into, rectifiers, surpassing, level, imagenet, 1502, 01852, kumar, siddharth, krishna, 1704, 08863, 256, 249, difficulty, bottou, leon, berlin, heidelberg, 8_2, tricks, trade, efficient, backprop, shin, yeonjong, yanhui, karniadakis, numerical, 1706, 4208, cicp, 0165, 1903, 06733, 1671, communications, massachusetts, jozefowicz, rafal, zaremba, wojciech, 2350, 2342, 32nd, exploration, navdeep, simple, rectified, 1504, 00941, tension, having, tradeoffs, causes, undesirable, trait, while, developed, automatically, tune, optimizers, generated, considerable, excitement, phase, possible, however, paper, demonstrated, well, chosen, hyperparameters, sufficient, needing, either, combination, 2010s, era, common, difficult, starting, bottom, contrastive, divergence, belief, early, described, perceptrons, attracting, fixed, recommend, randomly, begin, leq, particular, scaling, maps, interval, itself, thus, ensuring, overall, gain, around, operating, conditions, maximum, which, improves, hyperbolic, tangent, behave, any, let, concatenation, looks, norm, unbiased, moves, instead, subset, larger, total, transformers, multiplier, element, inside, branches, follows, introduction, allowed, deeper, than, previous, gave, rise, their, own, connection, could, stabilize, unnecessary, normalizations, vgg, been, generalized, fully, connected, orthonormal, proceeding, runs, divides, deviation, approximately, sequential, lsuv, proposes, parameterize, result, throughout, remain, found, sequence, modelling, odd, widths, heights, central, fill, illustration, filling, stride, padding, delta, iid, calculate, transpose, depending, whether, tall, wide, left, right, according, multiplied, depends, time, until, independent, depth, haar, measure, poorly, equal, identically, compromise, goals, continuous, popularized, activations, most, situations, multiplicative, signal, positive, value, likely, 305, typically, bounded, range, unbounded, cause, suggested, parts, identity, similar, idea, assume, unless, otherwise, stated, simplest, form, but, leads, causing, features, symmetry, kinds, where, number, discuss, specific, discussed, later, sections, note, even, titled, choice, affects, speed, within, signals, quality, final, proper, necessary, issues, saturation, creating, modified, assigning, outline, research, ijcai, iclr, neurips, iccv, emnlp, ecml, pkdd, eccv, aaai, journals, conferences, topological, statistical, pac, occam, risk, minimization, machines, mathematical, roc, confusion, coefficient, determination, diagnostics, mechanistic, interpretability, loop, crowdsourcing, active, humans, play, multi, temporal, difference, ecram, electrochemical, ram, memtransistor, spiking, informed, radiance, deepdream, lenet, som, restricted, boltzmann, reservoir, computing, esn, isolation, local, outlier, ransac, markov, conditional, graphical, sdl, sne, pgd, pca, nmf, lda, ica, cca, exploratory, shift, optics, dbscan, expectation, maximization, fuzzy, hierarchical, cure, birch, support, svm, relevance, rvm, logistic, naive, boosting, bagging, ensembles, decision, trees, apprenticeship, multimodal, ontology, grammar, induction, rank, association, rules, automl, cleaning, density, estimation, modeling, quantum, neuromorphic, curriculum, online, few, meta, transfer, paradigms, mining, part, series, remove, message, please, looking, more, unreliable, citations, challenged, removed, listed, technique, encyclopedia, item, printable, version, download, print, export, switch, legacy, parser, get, shortened, url, permanent, here, actions, english, talk, subsection, personal, special, pages, recent, community, portal, contribute, current, events, navigation, jump, content,
Text of the page (random words):
network architectures are discussed in later sections for an mlp there are only two kinds of trainable parameters called weights and biases each layer l displaystyle l contains a weight matrix w l r n l 1 n l displaystyle w l in mathbb r n_ l 1 times n_ l and a bias vector b l r n l displaystyle b l in mathbb r n_ l where n l displaystyle n_ l is the number of neurons in that layer a weight initialization method is an algorithm for setting the initial values for w l b l displaystyle w l b l for each layer l displaystyle l the simplest form is zero initialization w l 0 b l 0 displaystyle w l 0 b l 0 zero initialization is usually used for initializing biases but it is not used for initializing weights as it leads to symmetry in the network causing all neurons to learn the same features in this page we assume b 0 displaystyle b 0 unless otherwise stated recurrent neural networks typically use activation functions with bounded range such as sigmoid and tanh since unbounded activation may cause exploding values le jaitly hinton 2015 1 suggested initializing weights in the recurrent parts of the network to identity and zero bias similar to the idea of residual connections and lstm with no forget gate in most cases the biases are initialized to zero though some situations can use a nonzero initialization for example in multiplicative units such as the forget gate of lstm the bias can be initialized to 1 to allow good gradient signal through the gate 2 for neurons with relu activation one can initialize the bias to a small positive value like 0 1 so that the gradient is likely nonzero at initialization avoiding the dying relu problem 3 305 4 random initialization edit random initialization means sampling the weights from a normal distribution or a uniform distribution usually independently lecun initialization edit lecun initialization popularized in lecun et al 1998 5 is designed to preserve the variance of neural activations during the forward pass it samples each entry in w l displaystyle w l independently from a distribution with mean 0 and variance 1 n l 1 displaystyle 1 n_ l 1 for example if the distribution is a continuous uniform distribution then the distribution is u 3 n l 1 displaystyle mathcal u pm sqrt 3 n_ l 1 glorot initialization edit glorot initialization or xavier initialization was proposed by xavier glorot and yoshua bengio 6 it was designed as a compromise between two goals to preserve activation variance during the forward pass and to preserve gradient variance during the backward pass for uniform initialization it samples each entry in w l displaystyle w l independently and identically from u 6 n l 1 n l 1 displaystyle mathcal u pm sqrt 6 n_ l 1 n_ l 1 in the context n l 1 displaystyle n_ l 1 is also called the fan in and n l 1 displaystyle n_ l 1 the fan out when the fan in and fan out are equal then glorot initialization is the same as lecun initialization he initialization edit as glorot initialization performs poorly for relu activation 7 he initialization or kaiming initialization was proposed by kaiming he et al 8 for networks with relu activation it samples each entry in w l displaystyle w l from n 0 2 n l 1 displaystyle mathcal n 0 2 n_ l 1 orthogonal initialization edit saxe et al 2013 9 proposed orthogonal initialization initializing weight matrices as uniformly random according to the haar measure semi orthogonal matrices multiplied by a factor that depends on the activation function of the layer it was designed so that if one initializes a deep linear network this way then its training time until convergence is independent of depth 10 sampling a uniformly random semi orthogonal matrix can be done by initializing x displaystyle x by iid sampling its entries from a standard normal distribution then calculate x x 1 2 x displaystyle left xx top right 1 2 x or its transpose depending on whether x displaystyle x is tall or wide 11 for cnn kernels with odd widths and heights orthogonal initialization is done this way initialize the central point by a semi orthogonal matrix and fill the other entries with zero as an illustration a kernel k displaystyle k of shape 3 3 c c displaystyle 3 times 3 times c times c is initialized by filling k 2 2 displaystyle k 2 2 with the entries of a random semi orthogonal matrix of shape c c displaystyle c times c and the other entries with zero balduzzi et al 2017 12 used it with stride 1 and zero padding this is sometimes called the orthogonal delta initialization 11 13 related to this approach unitary initialization proposes to parameterize the weight matrices to be unitary matrices with the result that at initialization they are random unitary matrices and throughout training they remain unitary this is found to improve long sequence modelling in lstm 14 15 orthogonal initialization has been generalized to layer sequential unit variance lsuv initialization it is a data dependent initialization method and can be used in convolutional neural networks it first initializes weights of each convolution or fully connected layer with orthonormal matrices then proceeding from the first to the last layer it runs a forward pass on a random minibatch and divides the layer s weights by the standard deviation of its output so that its output has variance approximately 1 16 17 fixup initialization edit in 2015 the introduction of residual connections allowed very deep neural networks to be trained much deeper than the 20 layers of the previous state of the art such as the vgg 19 residual connections gave rise to their own weight initialization problems and strategies these are sometimes called normalization free methods since using residual connection could stabilize the training of a deep neural network so much that normalizations become unnecessary fixup initialization is designed specifically for networks with residual connections and without batch normalization as follows 18 initialize the classification layer and the last layer of each residual branch to 0 initialize every other layer using a standard method such as he initialization and scale only the weight layers inside residual branches by l 1 2 m 2 displaystyle l frac 1 2m 2 add a scalar multiplier initialized at 1 in every branch and a scalar bias initialized at 0 before each convolution linear and element wise activation layer similarly t fixup initialization is designed for transformers without layer normalization 19 9 others edit instead of initializing all weights with random values on the order of o 1 n displaystyle o 1 sqrt n sparse initialization initialized only a small subset of the weights with larger random values and the other weights zero so that the total variance is still on the order of o 1 displaystyle o 1 20 random walk initialization was designed for mlp so that during backpropagation the l2 norm of gradient at each layer performs an unbiased random walk as one moves from the last layer to the first 21 looks linear initialization was designed to allow the neural network to behave like a deep linear network at initialization since w r e l u x w r e l u x w x displaystyle w mathrm relu x w mathrm relu x wx it initializes a matrix w displaystyle w of shape r n 2 m displaystyle mathbb r frac n 2 times m by any method such as orthogonal initialization then let the r n m displaystyle mathbb r n times m weight matrix to be the concatenation of w w displaystyle w w 22 miscellaneous edit for hyperbolic tangent activation function a particular scaling is sometimes used 1 7159 tanh 2 x 3 displaystyle 1 7159 tanh 2x 3 this was sometimes called lecun s tanh it was designed so that it maps the interval 1 1 displaystyle 1 1 to itself thus ensuring that the overall gain is around 1 in normal operating conditions and that f x displaystyle f x is at maximum when x 1 1 displaystyle x 1 1 which improves convergence at the end of training 23 5 in self normalizing neural networks the selu activation function s e l u x λ x if x 0 α e x α if x 0 displaystyle mathrm selu x lambda begin cases x text if x 0 alpha e x alpha text if x leq 0 end cases with parameters λ 1 0507 α 1 6733 displaystyle lambda approx 1 0507 alpha approx 1 6733 makes it such that the mean and variance of the output of each layer has 0 1 displaystyle 0 1 as an attracting fixed point this makes initialization less important though they recommend initializing weights randomly with variance 1 n l 1 displaystyle 1 n_ l 1 24 history edit random weight initialization was used since frank rosenblatt s perceptrons an early work that described weight initialization specifically was lecun et al 1998 5 before the 2010s era of deep learning it was common to initialize models by generative pre training using an unsupervised learning algorithm that is not backpropagation as it was difficult to directly train deep neural networks by backpropagation 25 26 for example a deep belief network was trained by using contrastive divergence layer by layer starting from the bottom 27 martens 2010 20 proposed hessian free optimization a quasi newton method to directly train deep networks the work generated considerable excitement that initializing networks without pre training phase was possible 28 however a 2013 paper demonstrated that with well chosen hyperparameters momentum gradient descent with weight initialization was sufficient for training neural networks without needing either quasi newton method or generative pre training a combination that is still in use as of 2024 29 since then the impact of initialization on tuning the variance has become less important with methods developed to automatically tune variance like batch normalization tuning the variance of the forward pass 30 and momentum based optimizers tuning the variance of the backward pass 31 there is a tension between using careful weight initialization to decrease the need for normalization and using normalization to decrease the need for careful weight initialization with each approach having its tradeoffs for example batch normalization causes training examples in the minibatch to become dependent an undesirable trait while weight initialization is architecture dependent 32 see also edit backpropagation normalization machine learning gradient descent vanishing gradient problem references edit le quoc v jaitly navdeep hinton geoffrey e 2015 a simple way to initialize recurrent networks of rectified linear units arxiv 1504 00941 cs ne jozefowicz rafal zaremba wojciech sutskever ilya 2015 06 01 an empirical exploration of recurrent network architectures proceedings of the 32nd international conference on machine learning pmlr 2342 2350 goodfellow ian bengio yoshua courville aaron 2016 deep learning adaptive computation and machine learning cambridge massachusetts the mit press isbn 978 0 262 03561 3 lu lu shin yeonjong su yanhui karniadakis george em 2019 dying relu and initialization theory and numerical examples communications in computational physics 28 5 1671 1706 arxiv 1903 06733 doi 10 4208 cicp oa 2020 0165 1 2 3 lecun yann bottou leon orr genevieve b müller klaus robert 1998 efficient backprop in orr genevieve b müller klaus robert eds neural networks tricks of the trade berlin heidelberg springer pp 9 50 doi 10 1007 3 540 49430 8_2 isbn 978 3 540 49430 0 retrieved 2024 10 05 glorot xavier bengio yoshua 2010 03 31 understanding the difficulty of training deep feedforward neural networks proceedings of the thirteenth international conference on artificial intelligence and statistics jmlr workshop and conference proceedings 249 256 kumar siddharth krishna 2017 on weight initialization in deep neural networks arxiv 1704 08863 cs lg he kaiming zhang xiangyu ren shaoqing sun jian 2015 delving deep into rectifiers surpassing human level performance on imagenet classification arxiv 1502 01852 cs cv saxe andrew m mcclelland james l ganguli surya 2013 exact solutions to the nonlinear dynamics of learning in deep linear neural networks arxiv 1312 6120 cs ne hu wei xiao lechao pennington jeffrey 2020 provable benefit of orthogonal initialization in optimizing deep linear networks arxiv 2001 05992 cs lg 1 2 martens james ballard andy desjardins guillaume swirszcz grzegorz dalibard valentin sohl dickstein jascha schoenholz samuel s 2021 rapid training of deep neural networks without skip connections or normalization layers using deep kernel shaping arxiv 2110 01765 cs lg balduzzi david frean marcus leary lennox lewis j p ma kurt wan duo mcwilliams brian 2017 07 17 the shattered gradients problem if resnets are the answer then what is the question proceedings of the 34th international conference on machine learning pmlr 342 350 xiao lechao bahri yasaman sohl dickstein jascha schoenholz samuel pennington jeffrey 2018 07 03 dynamical isometry and a mean field theory of cnns how to train 10 000 layer vanilla convolutional neural networks proceedings of the 35th international conference on machine learning pmlr 5393 5402 arxiv 1806 05393 arjovsky martin shah amar bengio yoshua 2016 06 11 unitary evolution recurrent neural networks proceedings of the 33rd international conference on machine learning pmlr 1120 1128 arxiv 1511 06464 henaff mikael szlam arthur lecun yann 2017 03 15 recurrent orthogonal networks and long memory tasks arxiv 1602 06662 cs ne mishkin dmytro matas jiri 2016 02 19 all you need is a good init arxiv 1511 06422 xie di xiong jiang pu shiliang 2017 all you need is beyond a good init exploring better solution for training extremely deep convolutional neural networks with orthonormality and modulation ieee conference on computer vision and pattern recognition cvpr pp 6176 6185 zhang hongyi dauphin yann n ma tengyu 2019 fixup initialization residual learning without normalization arxiv 1901 09321 cs lg huang xiao shi perez felipe ba jimmy volkovs maksims 2020 11 21 improving transformer optimization through better initialization proceedings of the 37th international conference on machine learning pmlr 4475 4483 1 2 martens james 2010 06 21 deep learning via hessian free optimization proceedings of the 27th international conference on international conference on machine learning icml 10 madison wi usa omnipress 735 742 isbn 978 1 60558 907 7 cite journal cs1 maint periodical has isbn link sussillo david abbott l f 2014 random walk initialization for training very deep feedforward networks arxiv 1412 6558 cs ne balduzzi david frean marcus leary lennox lewis jp kurt wan duo ma mcwilliams brian 2017 the shattered gradients problem if resnets are the answer then what is the question arxiv 1702 08591 cs ne lecun y 1989 generalization and network design strategies pdf in pfeifer r schreter z fogelman f steels l eds connectionism in perspective proceedings of the international conference connectionism in perspective university of zurich 10 13 october 1988 amsterdam elsevier klambauer günter unterthiner thomas mayr andreas hochreiter sepp 2017 self normalizing neural networks advances in neural information...
|