Meta tags:
Headings (most frequently used words):
autoencoder, training, cae, contents, mathematical, principles, variations, advantages, of, depth, history, applications, see, also, further, reading, references, definition, an, interpretation, variational, vae, sparse, sae, denoising, dae, contractive, minimum, description, length, mdl, ae, concrete, dimensionality, reduction, information, retrieval, anomaly, detection, image, processing, drug, discovery, popularity, prediction, machine, translation, communication, systems, principal, component, analysis,
Text of the page (most frequently used words):
the (319), displaystyle (119), and (87), #autoencoder (86), learning (81), for (68), noise (64), autoencoders (52), with (46), neural (42), that (40), data (38), phi (34), deep (33), doi (33), edit (31), are (31), machine (29), theta (29), code (27), sparse (26), network (25), mathcal (25), training (24), rho (24), this (23), input (23), from (22), reconstruction (22), can (22), text (21), image (21), denoising (20), layer (20), networks (19), loss (18), arxiv (18), principal (18), two (18), function (18), systems (17), latent (17), s2cid (17), detection (17), which (17), hidden (16), isbn (16), anomaly (16), message (16), sparsity (16), representation (15), information (15), distribution (15), hinton (14), model (14), encoder (14), 978 (14), dimensionality (14), reduction (13), variational (13), 2017 (13), where (13), space (13), models (12), decoder (12), one (12), auto (12), small (12), using (11), international (11), 2019 (11), pmid (11), bibcode (11), used (11), set (11), then (11), mathbb (11), hat (11), error (10), applications (10), processing (10), analysis (10), component (10), learn (10), pca (10), but (10), min (10), dae (9), artificial (9), regularization (9), ieee (9), conference (9), 1016 (9), components (9), have (9), not (9), features (9), its (9), ref (9), was (8), has (8), description (8), communication (8), ratio (8), engineering (8), gradient (8), images (8), supervised (8), 2014 (8), 2015 (8), single (8), linear (8), science (8), encoding (8), feature (8), length (8), contractive (8), such (8), applied (8), usually (8), first (8), most (8), empirical (8), main (8), defined (8), search (7), wikipedia (7), use (7), means (7), general (7), signal (7), generative (7), self (7), recognition (7), activation (7), functions (7), translation (7), 1109 (7), journal (7), issn (7), encoders (7), mathematical (7), how (7), also (7), more (7), while (7), sim (7), entries (7), same (7), optimal (7), vector (7), toggle (6), diffusion (6), gaussian (6), related (6), density (6), bengio (6), based (6), language (6), into (6), than (6), 2016 (6), boltzmann (6), stat (6), representations (6), probability (6), case (6), salakhutdinov (6), retrieval (6), concrete (6), minimum (6), useful (6), other (6), there (6), only (6), quantile (6), frac (6), classification (6), each (6), depth (6), cae (6), mdl (6), norm (6), much (6), any (6), cont (6), messages (6), sum (6), log (6), may (5), page (5), unsupervised (5), variation (5), value (5), vae (5), convolutional (5), transformer (5), yoshua (5), geoffrey (5), reasoning (5), speech (5), regression (5), history (5), 2018 (5), 2013 (5), machines (5), they (5), proceedings (5), press (5), 2009 (5), 313 (5), during (5), subsection (5), performance (5), anomalies (5), however (5), some (5), compared (5), many (5), right (5), efficient (5), subspace (5), dataset (5), advantages (5), possible (5), standard (5), delta (5), randomly (5), given (5), task (5), table (4), contents (4), policy (4), terms (4), dimension (4), architectures (4), local (4), low (4), methods (4), list (4), imaging (4), human (4), optimization (4), principle (4), vision (4), john (4), david (4), knowledge (4), cognitive (4), semantic (4), algorithm (4), backpropagation (4), method (4), cho (4), popularity (4), modeling (4), simple (4), end (4), compression (4), what (4), study (4), acm (4), mit (4), belief (4), 2006 (4), reducing (4), 1007 (4), decomposition (4), zipser (4), 1987 (4), example (4), foundations (4), examples (4), joint (4), better (4), 2011 (4), kramer (4), basic (4), liou (4), 031 (4), been (4), traditional (4), another (4), prediction (4), found (4), all (4), binary (4), perfect (4), learned (4), dimensional (4), typically (4), threshold (4), left (4), would (4), their (4), size (4), units (4), tasks (4), layers (4), variations (4), learns (4), reference (4), principles (4), robustness (4), nabla (4), measures (4), enforce (4), zero (4), lambda (4), way (4), quality (4), hide (4), move (4), sidebar (4), add (3), view (3), non (3), inc (3), short (3), wikidata (3), filter (3), interference (3), distortion (3), energy (3), audio (3), equivalent (3), effective (3), level (3), shot (3), video (3), health (3), architecture (3), gan (3), multilayer (3), recurrent (3), computer (3), noam (3), ian (3), goodfellow (3), alex (3), james (3), selection (3), world (3), stable (3), facial (3), synthesis (3), intelligence (3), agent (3), context (3), nmt (3), automated (3), approach (3), symbolic (3), reinforcement (3), datasets (3), weight (3), descent (3), clustering (3), bias (3), variance (3), algorithms (3), structure (3), jun (3), january (3), medical (3), stacked (3), december (3), mining (3), advances (3), 1145 (3), discovery (3), robust (3), special (3), nonlinear (3), baldi (3), computational (3), 2008 (3), fashion (3), mnist (3), bayes (3), 1988 (3), association (3), singular (3), elman (3), framework (3), without (3), autoassociative (3), neucom (3), neurocomputing (3), will (3), pdf (3), cheng (3), further (3), optimize (3), output (3), encoded (3), results (3), were (3), trained (3), application (3), setting (3), labels (3), exist (3), larger (3), good (3), anomalous (3), perform (3), points (3), normal (3), train (3), could (3), after (3), original (3), sigma (3), indeed (3), vectors (3), similar (3), activations (3), weights (3), span (3), dimensions (3), called (3), identity (3), pair (3), restricted (3), technique (3), objective (3), depends (3), over (3), theory (3), limit (3), daes (3), make (3), perturbations (3), caes (3), maps (3), desired (3), matrix (3), process (3), problem (3), version (3), schema (3), simply (3), define (3), sae (3), close (3), article (3), ideal (3), generate (3), just (3), parametrized (3), field (3), random (3), tools (3), languages (2), safety (2), contact (2), about (2), privacy (2), under (2), last (2), september (2), 2026 (2), categories (2), cs1 (2), maint (2), periodical (2), articles (2), total (2), pass (2), thermal (2), analyzer (2), topics (2), contrast (2), quantization (2), modulation (2), per (2), carrier (2), statistical (2), phase (2), temperature (2), channel (2), white (2), additive (2), masking (2), government (2), regulation (2), control (2), physics (2), impact (2), visual (2), social (2), act (2), mamba (2), rnn (2), perceptron (2), mlp (2), gru (2), memory (2), lstm (2), differentiable (2), turing (2), quoc (2), oriol (2), vinyals (2), ilya (2), sutskever (2), krizhevsky (2), fei (2), andrew (2), jürgen (2), schmidhuber (2), paul (2), rule (2), classifiers (2), ibm (2), large (2), generation (2), dall (2), alexnet (2), protocol (2), neuro (2), open (2), source (2), coding (2), word (2), rlhf (2), sarsa (2), sigmoid (2), tradeoff (2), projects (2), glossary (2), review (2), chinese (2), deeper (2), sequence (2), 1409 (2), kyunghyun (2), properties (2), approaches (2), predicting (2), posts (2), 5090 (2), cscita (2), 2nd (2), computing (2), designed (2), 1040 (2), 1038 (2), 2020 (2), hdl (2), biomedical (2), informatics (2), alzheimer (2), disease (2), liu (2), transactions (2), breast (2), cancer (2), song (2), various (2), isbi (2), zeng (2), wang (2), super (2), resolution (2), cybernetics (2), icdmw (2), new (2), improves (2), corrupted (2), lossy (2), score (2), know (2), 105 (2), zhou (2), 4503 (2), lecture (2), australia (2), semi (2), diagnosis (2), macao (2), plaut (2), subspaces (2), 1804 (2), 10253 (2), pierre (2), ontology (2), biology (2), graphical (2), hashing (2), adaptive (2), boosting (2), torch (2), diederik (2), kingma (2), welling (2), max (2), 1312 (2), 5947 (2), scholarpedia (2), 5786 (2), 507 (2), 1658773 (2), 16873662 (2), 1126 (2), 1127647 (2), 2006sci (2), 504h (2), 504 (2), perceptrons (2), cottrell (2), munro (2), society (2), internal (2), 0001 (2), harrison (2), continuous (2), cambridge (2), 1986 (2), parallel (2), distributed (2), 262 (2), hornik (2), 1989 (2), 0893 (2), 6080 (2), oja (2), 1982 (2), 273 (2), neuron (2), free (2), icml (2), chemical (2), 0098 (2), 1354 (2), link (2), cite (2), our (2), 1561 (2), trends (2), mark (2), 1991 (2), yuan (2), wei (2), words (2), media (2), springer (2), criterion (2), research (2), courville (2), aaron (2), bank (2), dor (2), koenigstein (2), giryes (2), raja (2), 2023 (2), 24627 (2), 24628 (2), 9_16 (2), handbook (2), references (2), reading (2), see (2), because (2), help (2), errors (2), inherent (2), difficulty (2), accurately (2), behavior (2), real (2), referred (2), sequences (2), procedure (2), generated (2), specific (2), still (2), online (2), strategies (2), experimentally (2), drug (2), experiments (2), between (2), characteristics (2), best (2), considered (2), rare (2), events (2), very (2), these (2), recent (2), certain (2), autoencoding (2), reconstructing (2), understood (2), those (2), when (2), should (2), cases (2), reconstruct (2), validation (2), since (2), become (2), spaces (2), mapping (2), support (2), query (2), less (2), powerful (2), lower (2), solution (2), orthogonal (2), projection (2), equal (2), multi (2), smaller (2), both (2), identical (2), cost (2), early (2), discrete (2), like (2), etc (2), developed (2), module (2), generalized (2), showed (2), layered (2), exponentially (2), number (2), through (2), subset (2), minimize (2), infinitesimal (2), resist (2), encourage (2), obtained (2), aim (2), frobenius (2), jacobian (2), want (2), itself (2), expected (2), thus (2), neighborhood (2), even (2), fraction (2), chosen (2), include (2), type (2), noisy (2), circ (2), inputs (2), giving (2), cdot (2), upon (2), let (2), need (2), measure (2), times (2), infty (2), instead (2), following (2), top (2), yellow (2), variants (2), codes (2), fewer (2), active (2), vaes (2), families (2), different (2), receives (2), capacity (2), overcomplete (2), undercomplete (2), really (2), appears (2), parts (2), reconstructs (2), interpretation (2), euclidean (2), arg (2), write (2), refer (2), decoded (2), family (2), rightarrow (2), definition (2), subsequent (2), problems (2), curve (2), net (2), forest (2), factor (2), structured (2), appearance (2), upload (2), file (2), changes (2), links (2), read (2), create (2), account (2), donate (2), menu (2), topic, mobile, cookie, statement, statistics, developers, conduct, legal, contacts, disclaimers, available, additional, apply, site, you, agree, registered, trademark, profit, organization, wikimedia, foundation, creative, commons, attribution, sharealike, license, rendered, parsoid, edited, utc, dmy, dates, matches, retrieved, https, org, index, php, title, oldid, 1375973936, prior, shrinkage, fields, bm3d, block, matching, filtering, bilateral, anisotropic, blur, wavelet, median, denoise, radiation, spectrum, generator, colors, acoustics, cnr, sqnr, sinr, plus, snr, sinad, mer, symbol, bit, dbrnc, receiver, ratios, pseudorandom, nvh, vibration, harshness, spectral, shaping, floor, figure, impulse, pulse, resistance, circuit, coherent, worley, pink, johnson, nyquist, jitter, infrasound, grey, flicker, cosmic, burst, brownian, background, atmospheric, awgn, class, transportation, sound, ships, rooms, radio, environment, electronics, buildings, power, measurement, cancellation, acoustic, quieting, telecommunications, category, workplace, warfare, military, art, games, marketing, chatbot, psychosis, healthcare, fiction, education, engine, explainable, environmental, competition, arms, race, anthropomorphism, winter, slop, literacy, infrastructure, effect, center, bubble, boom, economic, opposition, centers, propaganda, virtual, politician, precautionary, nationalism, ethics, elections, takeover, alignment, cold, war, political, graph, gnn, adversarial, highway, residual, cnn, echo, state, gated, unit, long, term, vit, françois, chollet, daniel, kokotajlo, jan, leike, mustafa, suleyman, schulman, aidan, gomez, shazeer, ashish, vaswani, andrej, karpathy, silver, demis, hassabis, goodnight, graves, stephen, grossberg, lotfi, zadeh, yann, lecun, hopfield, werbos, seppo, linnainmaa, seymour, papert, joseph, weizenbaum, bernard, widrow, frank, rosenblatt, oliver, selfridge, herbert, simon, cliff, shaw, allen, newell, nathaniel, rochester, mccarthy, marvin, minsky, takeo, kanade, kunihiko, fukushima, shun, ichi, amari, claude, shannon, christopher, manning, von, neumann, walter, pitts, warren, sturgis, mcculloch, alan, people, yago, dbpedia, conceptnet, bases, opencog, lida, clarion, soar, reasoners, procedural, logic, programs, inference, engines, expert, deductive, robot, autogpt, action, muzero, driving, car, openai, five, alphazero, alphago, decisional, watsonx, watson, project, debater, oasis, genie, udio, suno, riffusion, music, veo, seedance, sora, kling, hailuo, runway, gen, dream, recraft, midjourney, imagen, ideogram, gpt, flux, firefly, aurora, alphafold, whisper, elevenlabs, ocr, hwr, wavenet, implementations, physical, agent2agent, hypothetical, superintelligence, asi, agi, weak, lethal, autonomous, weapons, laws, humanity, exam, companion, intelligent, game, playing, theorem, proving, actor, critic, situated, sovereign, blended, vibe, embedding, hallucination, recursive, improvement, reflection, llm, post, uncanny, valley, rag, adversary, autoregression, imitation, prompt, augmentation, initialization, gating, rectifier, softmax, batchnorm, normalization, convolution, attention, conjugate, quasi, newton, sgd, overfitting, double, hyperparameter, parameter, constraint, satisfaction, planning, concepts, lists, software, proprietary, institutions, companies, timeline, alnaseri, omar, alzubaidi, laith, himeur, yassine, timmermann, jens, 2024, design, next, eess, 2412, 13843, han, lifeng, kuang, shaohui, incorporating, radicals, character, 1805, 01565, 3215, bart, van, merrienboer, bahdanau, dzmitry, 1259, shaunak, maity, abhishek, goel, vritti, shitole, sanjay, bhattacharya, avik, instagram, lifestyle, magazine, 177, 35350962, 4381, 8066548, 174, gregory, barber, wired, molecule, exhibits, druglike, qualities, zhavoronkov, enables, rapid, identification, potent, ddr1, kinase, inhibitors, 201716327, 31477924, s41587, 019, 0224, nature, biotechnology, martinez, murcia, francisco, ortiz, andres, gorriz, juan, ramirez, javier, castillo, barnes, diego, 195187846, 31217131, 10630, 28806, jbhi, 2914970, 2020ijbhi, 17m, studying, manifold, xiang, lei, qingshan, gilmore, hannah, jianzhong, tang, jinghai, madabhushi, anant, 130, 26208307, 4729702, pmc, tmi, 2458702, 2016itmi, 119x, 119, ssae, nuclei, histopathology, tzu, hsi, sanchez, victor, hesham, eidaly, nasir, rajpoot, hybrid, curvature, types, cells, bone, marrow, trephine, biopsy, 1043, 7433130, 1172, 7950694, 14th, symposium, kun, ruxin, cuihua, tao, dacheng, coupled, 20787612, 26625442, 2168, 2267, tcyb, 2501373, 2017itcyb, 27z, gondara, lovedeep, barcelona, spain, 246, 14354973, 9781509059102, 0041, 2016arxiv160804667g, 1608, 04667, 241, 16th, workshops, buades, coll, morel, 2005, 530, 218466166, 1137, 040616024, 490, multiscale, simulation, 1301, 3468, february, sparsification, highly, 432, 440, balle, laparra, simoncelli, april, optimized, 1611, 01704, theis, lucas, shi, wenzhe, cunningham, huszár, ferenc, compressive, 1703, 00395, xiao, zhisheng, yan, qing, amit, yali, 2003, 02977, likelihood, regret, out, nalisnick, eric, matsukawa, akihiro, teh, yee, whye, gorur, dilan, lakshminarayanan, balaji, don, 1810, 09136, ribeiro, manassés, lazzaretti, andré, eugênio, lopes, heitor, silvério, videos, patrec, 016, 2018parel, 13r, pattern, letters, chong, paffenroth, randy, 674, 207557733, 4887, 3097983, 3098052, 665, 23rd, sigkdd, sakurada, mayu, yairi, takehisa, gold, coast, qld, 14613395, 3159, 2689746, 2689747, mlsda, workshop, sensory, morales, forero, bassetto, methodology, 1037, 211027131, 7281, 3804, ieem44572, 8978509, 1031, industrial, management, ieem, chicco, davide, sadowski, peter, gene, annotation, predictions, 533, 207217210, 9781450328944, 11311, 964622, 2649387, 2649442, 5th, bioinformatics, bcb, ruslan, section, 0888, 613x, ijar, 006, 969, approximate, github, schwenk, holger, 1997, ackley, sejnowski, march, 1985, 169, s0364, 0213, 80012, 147, generating, faces, boesen, larsen, sonderby, blog, html, 6114, overview, 117, 11715509, 25462637, neunet, 003, 1404, 7828, 4249, 2009schpj, 5947h, bourlard, kamp, 294, 206775335, 3196773, bf00332918, 291, biological, garrison, annual, meeting, gray, scale, extensional, programming, jeffrey, 1626, 3372872, 4966, 1121, 395916, 1988asaj, 1615e, 1615, acoustical, america, connectionist, university, dissertation, rumelhart, mcclelland, 29140, 7551, mitpress, 5236, 001, explorations, microstructure, cognition, kurt, 90014, minima, erkki, 7153672, 1432, 1416, bf00275687, 267, simplified, aistats, 448, 455, yingbo, arpit, devansh, nwogu, ifeoma, govindaraju, venu, 1405, 1380, elad, abid, abubakar, balin, muhammad, fatih, zou, 1901, 09346, zemel, richard, 1993, morgan, kaufmann, helmholtz, rifai, explicit, invariance, extraction, 1992, neutral, 328, 80051, computers, nianyin, zhang, hong, baoye, weibo, yurong, dobaie, abdullah, expression, via, 649, 0925, 2312, 043, 643, nair, vinod, nips, usa, curran, associates, 1347, 9781615679119, 1339, 22nd, object, nets, cs294a, notes, makhzani, alireza, frey, brendan, 5663, books, brain, 046506192, master, quest, ultimate, remake, domingos, pedro, 207178999, 23946944, 2200000006, 1795, 243, 1002, aic, 690370209, 1991aiche, 233k, 233, aiche, chen, jiun, daw, ran, 055, 139, huang, jau, chi, yang, wen, chie, 3150, 030, perception, géron, aurélien, canada, reilly, 740, 739, hands, scikit, keras, tensorflow, berlin, heidelberg, transforming, introduction, 392, 174802445, 2200000056, 2019arxiv190602691k, 1906, 02691, 307, vincent, pascal, larochelle, hugo, 2010, 3408, 3371, 0262035613, rokach, lior, maimon, oded, shmueli, erez, eds, 374, 353, computation, mass, 03561, cham, publishing, dictionary, important, resilient, impairments, crucial, transmitting, minimizing, addition, solve, several, limitations, designing, complex, channels, unlike, does, match, texts, treated, side, target, incorporate, rarely, done, due, availability, linguistic, recently, produced, promising, helpful, advertising, molecules, validated, mice, demanding, contexts, well, assisted, modelling, relation, decline, mri, preprocessing, outperformed, proved, competitive, against, jpeg, 2000, analyze, flagged, true, sense, metrics, fundamental, challenge, comes, gathered, imbalanced, indicating, introducing, estimates, confidence, intervals, evaluation, literature, shown, counterintuitively, consequently, able, reliably, intuitively, considering, rein, reconstructions, far, away, region, lie, axis, replicate, salient, constraints, described, previously, encouraged, precisely, reproduce, frequently, observed, facing, worsen, instances, others, frequency, observation, contribution, ignored, failing, unfamiliar, detect, recorded, percentile, taken, flag, estimate, correctly, asymptotically, grows, extreme, potentially, big, uncertainty, choice, estimated, implies, benefits, particularly, kinds, proposed, 2007, produce, database, stored, returning, slightly, flipping, bits, hash, potential, resides, linearity, allowing, generalizations, significantly, strongly, spanned, onto, generally, yet, recovered, them, 28x28pixel, come, improve, hallmark, place, semantically, near, pretrained, stack, initialize, gradually, until, hitting, bottleneck, neurons, resulting, yielded, qualitatively, easier, interpret, clearly, separating, clusters, rbms, plot, being, apart, rotation, selects, orientation, reflections, invariant, rotations, modern, associative, days, terminology, uncertain, associating, diabolo, date, 1990s, concept, became, widely, 2010s, involved, modules, generators, ais, immediately, resurgence, 1980s, suggested, put, mode, implemented, pairs, noted, subsequently, termed, researchers, debated, whether, whole, together, global, along, representative, layerwise, success, heavily, adopted, his, involves, treating, neighboring, pretraining, approximates, fine, tune, allows, show, whose, eigenvectors, yield, shallow, decrease, amount, needed, reduce, representing, often, decoders, offers, schematic, fully, connected, forces, consist, user, specified, uses, allow, gradients, selector, makes, categorical, relaxation, seeks, includes, expressed, represents, compressed, denotes, advanced, leverages, specifically, posits, provides, shortest, combined, ensure, compact, interpretable, finite, sized, extracted, additionally, stochastically, having, versions, point, analytically, penalizing, evaluated, samples, adds, ness, square, respect, understand, note, fact, property, leads, perhaps, pictures, look, exactly
Text of the page (random words):
odes that is e ϕ x displaystyle e_ phi x is close to zero in most entries sparse autoencoders may include more rather than fewer hidden units than inputs but only a small number of the hidden units are allowed to be active at the same time 12 encouraging sparsity improves performance on classification tasks 13 simple schema of a single layer sparse autoencoder the hidden nodes in bright yellow are activated while the light yellow ones are inactive the activation depends on the input there are two main ways to enforce sparsity one way is to simply clamp all but the highest k activations of the latent code to zero this is the k sparse autoencoder 13 the k sparse autoencoder inserts the following k sparse function in the latent layer of a standard autoencoder f k x 1 x n x 1 b 1 x n b n displaystyle f_ k x_ 1 x_ n x_ 1 b_ 1 x_ n b_ n where b i 1 displaystyle b_ i 1 if x i displaystyle x_ i ranks in the top k and 0 otherwise backpropagating through f k displaystyle f_ k is simple set gradient to 0 for b i 0 displaystyle b_ i 0 entries and keep gradient for b i 1 displaystyle b_ i 1 entries this is essentially a generalized relu function 13 the other way is a relaxed version of the k sparse autoencoder instead of forcing sparsity we add a sparsity regularization loss then optimize for min θ ϕ l θ ϕ λ l sparse θ ϕ displaystyle min _ theta phi l theta phi lambda l_ text sparse theta phi where λ 0 displaystyle lambda 0 measures how much sparsity we want to enforce 14 let the autoencoder architecture have k displaystyle k layers to define a sparsity regularization loss we need a desired sparsity ρ k displaystyle hat rho _ k for each layer a weight w k displaystyle w_ k for how much to enforce each sparsity and a function s 0 1 0 1 0 displaystyle s 0 1 times 0 1 to 0 infty to measure how much two sparsities differ for each input x displaystyle x let the actual sparsity of activation in each layer k displaystyle k be ρ k x 1 n i 1 n a k i x displaystyle rho _ k x frac 1 n sum _ i 1 n a_ k i x where a k i x displaystyle a_ k i x is the activation in the i displaystyle i th neuron of the k displaystyle k th layer upon input x displaystyle x the sparsity loss upon input x displaystyle x for one layer is s ρ k ρ k x displaystyle s hat rho _ k rho _ k x and the sparsity regularization loss for the entire autoencoder is the expected weighted sum of sparsity losses l sparse θ ϕ e x μ x k 1 k w k s ρ k ρ k x displaystyle l_ text sparse theta phi mathbb mathbb e _ x sim mu _ x left sum _ k in 1 k w_ k s hat rho _ k rho _ k x right typically the function s displaystyle s is either the kullback leibler kl divergence as 13 14 15 16 s ρ ρ k l ρ ρ ρ log ρ ρ 1 ρ log 1 ρ 1 ρ displaystyle s rho hat rho kl rho hat rho rho log frac rho hat rho 1 rho log frac 1 rho 1 hat rho or the l1 loss as s ρ ρ ρ ρ displaystyle s rho hat rho rho hat rho or the l2 loss as s ρ ρ ρ ρ 2 displaystyle s rho hat rho rho hat rho 2 alternatively the sparsity regularization loss may be defined without reference to any desired sparsity but simply force as much sparsity as possible in this case one can define the sparsity regularization loss as l sparse θ ϕ e x μ x k 1 k w k h k displaystyle l_ text sparse theta phi mathbb mathbb e _ x sim mu _ x left sum _ k in 1 k w_ k h_ k right where h k displaystyle h_ k is the activation vector in the k displaystyle k th layer of the autoencoder the norm displaystyle cdot is usually the l1 norm giving the l1 sparse autoencoder or the l2 norm giving the l2 sparse autoencoder denoising autoencoder dae edit a schema of a denoising autoencoder denoising autoencoders dae try to achieve a good representation by changing the reconstruction criterion 2 3 a dae originally called a robust autoassociative network by mark a kramer 17 is trained by intentionally corrupting the inputs of a standard autoencoder during training a noise process is defined by a probability distribution μ t displaystyle mu _ t over functions t x x displaystyle t mathcal x to mathcal x that is the function t displaystyle t takes a message x x displaystyle x in mathcal x and corrupts it to a noisy version t x displaystyle t x the function t displaystyle t is selected randomly with a probability distribution μ t displaystyle mu _ t given a task μ ref d displaystyle mu _ text ref d the problem of training a dae is the optimization problem min θ ϕ l θ ϕ e x μ x t μ t d x d θ e ϕ t x displaystyle min _ theta phi l theta phi mathbb mathbb e _ x sim mu _ x t sim mu _ t d x d_ theta circ e_ phi circ t x that is the optimal dae should take any noisy message and attempt to recover the original message without noise thus the name denoising usually the noise process t displaystyle t is applied only during training and testing not during downstream use the use of dae depends on two assumptions there exist representations to the messages that are relatively stable and robust to the type of noise we are likely to encounter the said representations capture structures in the input distribution that are useful for our purposes 3 example noise processes include additive isotropic gaussian noise masking noise a fraction of the input is randomly chosen and set to 0 salt and pepper noise a fraction of the input is randomly chosen and randomly set to its minimum or maximum value 3 contractive autoencoder cae edit a contractive autoencoder cae adds the contractive regularization loss to the standard autoencoder loss min θ ϕ l θ ϕ λ l cont θ ϕ displaystyle min _ theta phi l theta phi lambda l_ text cont theta phi where λ 0 displaystyle lambda 0 measures how much contractive ness we want to enforce the contractive regularization loss itself is defined as the expected square of frobenius norm of the jacobian matrix of the encoder activations with respect to the input l cont θ ϕ e x μ r e f x e ϕ x f 2 displaystyle l_ text cont theta phi mathbb e _ x sim mu _ ref nabla _ x e_ phi x _ f 2 to understand what l cont displaystyle l_ text cont measures note the fact e ϕ x δ x e ϕ x 2 x e ϕ x f δ x 2 displaystyle e_ phi x delta x e_ phi x _ 2 leq nabla _ x e_ phi x _ f delta x _ 2 for any message x x displaystyle x in mathcal x and small variation δ x displaystyle delta x in it thus if x e ϕ x f 2 displaystyle nabla _ x e_ phi x _ f 2 is small it means that a small neighborhood of the message maps to a small neighborhood of its code this is a desired property as it means small variation in the message leads to small perhaps even zero variation in its code like how two pictures may look the same even if they are not exactly the same the cae can be understood as an infinitesimal limit of dae in the limit of small gaussian input noise daes make the reconstruction function resist small but finite sized input perturbations while caes make the extracted features resist infinitesimal input perturbations however while daes encourage robustness of reconstruction caes encourage robustness of representation additionally daes robustness is obtained stochastically by having various corrupted versions of a training point aim for an identical reconstruction in contrast caes robustness to small perturbations is obtained analytically by penalizing the frobenius norm of the encoder jacobian x e ϕ x f 2 displaystyle nabla _ x e_ phi x _ f 2 evaluated at the training samples 18 minimum description length autoencoder mdl ae edit a minimum description length autoencoder mdl ae is an advanced variation of the traditional autoencoder which leverages principles from information theory specifically the minimum description length mdl principle the mdl principle posits that the best model for a dataset is the one that provides the shortest combined encoding of the model and the data in the context of autoencoders this principle is applied to ensure that the learned representation is not only compact but also interpretable and efficient for reconstruction the mdl ae seeks to minimize the total description length of the data which includes the size of the latent representation code length and the error in reconstructing the original data the objective can be expressed as l code l error displaystyle l_ text code l_ text error where l code displaystyle l_ text code represents the length of the compressed latent representation and l error displaystyle l_ text error denotes the reconstruction error 19 concrete autoencoder cae edit the concrete autoencoder is designed for discrete feature selection 20 a concrete autoencoder forces the latent space to consist only of a user specified number of features the concrete autoencoder uses a continuous relaxation of the categorical distribution to allow gradients to pass through the feature selector layer which makes it possible to use standard backpropagation to learn an optimal subset of input features that minimize reconstruction loss advantages of depth edit schematic structure of an autoencoder with 3 fully connected hidden layers the code z or h for reference in the text is the most internal layer autoencoders are often trained with a single layer encoder and a single layer decoder but using many layered deep encoders and decoders offers many advantages 2 depth can exponentially reduce the computational cost of representing some functions depth can exponentially decrease the amount of training data needed to learn some functions experimentally deep autoencoders yield better compression compared to shallow or linear autoencoders 10 depth allows for advantages over traditional methods as one can show that after training single layer linear autoencoders have a latent space whose vectors span the same subspace as the eigenvectors found in principal component analysis 21 training edit geoffrey hinton developed the deep belief network technique for training many layered deep autoencoders his method involves treating each neighboring set of two layers as a restricted boltzmann machine so that pretraining approximates a good solution then using backpropagation to fine tune the results 10 researchers have debated whether joint training i e training the whole architecture together with a single global reconstruction objective to optimize would be better for deep auto encoders 22 a 2015 study showed that joint training learns better data models along with more representative features for classification as compared to the layerwise method 22 however their experiments showed that the success of joint training depends heavily on the regularization strategies adopted 22 23 history edit oja 1982 24 noted that pca is equivalent to a neural network with one hidden layer with identity activation function in the language of autoencoding the input to hidden module is the encoder and the hidden to output module is the decoder subsequently in baldi and hornik 1989 25 and kramer 1991 9 generalized pca to autoencoders a technique which they termed nonlinear pca immediately after the resurgence of neural networks in the 1980s it was suggested in 1986 26 that a neural network be put in auto association mode this was then implemented in harrison 1987 27 and elman zipser 1988 28 for speech and in cottrell munro zipser 1987 29 for images 30 in hinton salakhutdinov 2006 31 deep belief networks were developed these train a pair restricted boltzmann machines as encoder decoder pairs then train another pair on the latent representation of the first pair and so on 32 the first applications of ae date to early 1990s 2 33 19 their most traditional application was dimensionality reduction or feature learning but the concept became widely used for learning generative models of data 34 35 some of the most powerful ais in the 2010s involved autoencoder modules as a component of larger ai systems such as vae in stable diffusion discrete vae in transformer based image generators like dall e 1 etc during the early days when the terminology was uncertain the autoencoder has also been called identity mapping 25 9 auto associating 36 self supervised backpropagation 9 or diabolo network 37 11 applications edit the two main applications of autoencoders are dimensionality reduction and information retrieval or associative memory 2 but modern variations have been applied to other tasks dimensionality reduction edit plot of the first two principal components left and a two dimension hidden layer of a linear autoencoder right applied to the fashion mnist dataset 38 the two models being both linear learn to span the same subspace the projection of the data points is indeed identical apart from rotation of the subspace while pca selects a specific orientation up to reflections in the general case the cost function of a simple autoencoder is invariant to rotations of the latent space dimensionality reduction was one of the first deep learning applications 2 for hinton s 2006 study 10 he pretrained a multi layer autoencoder with a stack of rbms and then used their weights to initialize a deep autoencoder with gradually smaller hidden layers until hitting a bottleneck of 30 neurons the resulting 30 dimensions of the code yielded a smaller reconstruction error compared to the first 30 components of a principal component analysis pca and learned a representation that was qualitatively easier to interpret clearly separating data clusters 2 10 reducing dimensions can improve performance on tasks such as classification 2 indeed the hallmark of dimensionality reduction is to place semantically related examples near each other 39 principal component analysis edit reconstruction of 28x28pixel images by an autoencoder with a code size of two two units hidden layer and the reconstruction from the first two principal components of pca images come from the fashion mnist dataset 38 if linear activations are used or only a single sigmoid hidden layer then the optimal solution to an autoencoder is strongly related to principal component analysis pca 30 40 the weights of an autoencoder with a single hidden layer of size p displaystyle p where p displaystyle p is less than the size of the input span the same vector subspace as the one spanned by the first p displaystyle p principal components and the output of the autoencoder is an orthogonal projection onto this subspace the autoencoder weights are not equal to the principal components and are generally not orthogonal yet the principal components may be recovered from them using the singular value decomposition 41 however the potential of autoencoders resides in their non linearity allowing the model to learn more powerful generalizations compared to pca and to reconstruct the input with significantly lower information loss 10 information retrieval edit information retrieval benefits particularly from dimensionality reduction in that search can become more efficient in certain kinds of low dimensional spaces autoencoders were indeed applied to semantic hashing proposed by salakhutdinov and hinton in 2007 39 by training the algorithm to produce a low dimensional binary code all database entries could be sto...
|