Meta tags:
Headings (most frequently used words):
autoencoder, training, cae, contents, mathematical, principles, variations, advantages, of, depth, history, applications, see, also, further, reading, references, definition, an, interpretation, variational, vae, sparse, sae, denoising, dae, contractive, minimum, description, length, mdl, ae, concrete, dimensionality, reduction, information, retrieval, anomaly, detection, image, processing, drug, discovery, popularity, prediction, machine, translation, communication, systems, principal, component, analysis,
Text of the page (most frequently used words):
the (319), displaystyle (119), and (87), #autoencoder (86), learning (81), for (68), noise (64), autoencoders (52), with (46), neural (42), that (40), data (38), phi (34), deep (33), doi (33), edit (31), are (31), machine (29), theta (29), code (27), sparse (26), network (25), mathcal (25), training (24), rho (24), this (23), input (23), from (22), reconstruction (22), can (22), text (21), image (21), denoising (20), layer (20), networks (19), loss (18), arxiv (18), principal (18), two (18), function (18), systems (17), latent (17), s2cid (17), detection (17), which (17), hidden (16), isbn (16), anomaly (16), message (16), sparsity (16), representation (15), information (15), distribution (15), hinton (14), model (14), encoder (14), 978 (14), dimensionality (14), reduction (13), variational (13), 2017 (13), where (13), space (13), models (12), decoder (12), one (12), auto (12), small (12), using (11), international (11), 2019 (11), pmid (11), bibcode (11), used (11), set (11), then (11), mathbb (11), hat (11), error (10), applications (10), processing (10), analysis (10), component (10), learn (10), pca (10), but (10), min (10), dae (9), artificial (9), regularization (9), ieee (9), conference (9), 1016 (9), components (9), have (9), not (9), features (9), its (9), ref (9), was (8), has (8), description (8), communication (8), ratio (8), engineering (8), gradient (8), images (8), supervised (8), 2014 (8), 2015 (8), single (8), linear (8), science (8), encoding (8), feature (8), length (8), contractive (8), such (8), applied (8), usually (8), first (8), most (8), empirical (8), main (8), defined (8), search (7), wikipedia (7), use (7), means (7), general (7), signal (7), generative (7), self (7), recognition (7), activation (7), functions (7), translation (7), 1109 (7), journal (7), issn (7), encoders (7), mathematical (7), how (7), also (7), more (7), while (7), sim (7), entries (7), same (7), optimal (7), vector (7), toggle (6), diffusion (6), gaussian (6), related (6), density (6), bengio (6), based (6), language (6), into (6), than (6), 2016 (6), boltzmann (6), stat (6), representations (6), probability (6), case (6), salakhutdinov (6), retrieval (6), concrete (6), minimum (6), useful (6), other (6), there (6), only (6), quantile (6), frac (6), classification (6), each (6), depth (6), cae (6), mdl (6), norm (6), much (6), any (6), cont (6), messages (6), sum (6), log (6), may (5), page (5), unsupervised (5), variation (5), value (5), vae (5), convolutional (5), transformer (5), yoshua (5), geoffrey (5), reasoning (5), speech (5), regression (5), history (5), 2018 (5), 2013 (5), machines (5), they (5), proceedings (5), press (5), 2009 (5), 313 (5), during (5), subsection (5), performance (5), anomalies (5), however (5), some (5), compared (5), many (5), right (5), efficient (5), subspace (5), dataset (5), advantages (5), possible (5), standard (5), delta (5), randomly (5), given (5), task (5), table (4), contents (4), policy (4), terms (4), dimension (4), architectures (4), local (4), low (4), methods (4), list (4), imaging (4), human (4), optimization (4), principle (4), vision (4), john (4), david (4), knowledge (4), cognitive (4), semantic (4), algorithm (4), backpropagation (4), method (4), cho (4), popularity (4), modeling (4), simple (4), end (4), compression (4), what (4), study (4), acm (4), mit (4), belief (4), 2006 (4), reducing (4), 1007 (4), decomposition (4), zipser (4), 1987 (4), example (4), foundations (4), examples (4), joint (4), better (4), 2011 (4), kramer (4), basic (4), liou (4), 031 (4), been (4), traditional (4), another (4), prediction (4), found (4), all (4), binary (4), perfect (4), learned (4), dimensional (4), typically (4), threshold (4), left (4), would (4), their (4), size (4), units (4), tasks (4), layers (4), variations (4), learns (4), reference (4), principles (4), robustness (4), nabla (4), measures (4), enforce (4), zero (4), lambda (4), way (4), quality (4), hide (4), move (4), sidebar (4), add (3), view (3), non (3), inc (3), short (3), wikidata (3), filter (3), interference (3), distortion (3), energy (3), audio (3), equivalent (3), effective (3), level (3), shot (3), video (3), health (3), architecture (3), gan (3), multilayer (3), recurrent (3), computer (3), noam (3), ian (3), goodfellow (3), alex (3), james (3), selection (3), world (3), stable (3), facial (3), synthesis (3), intelligence (3), agent (3), context (3), nmt (3), automated (3), approach (3), symbolic (3), reinforcement (3), datasets (3), weight (3), descent (3), clustering (3), bias (3), variance (3), algorithms (3), structure (3), jun (3), january (3), medical (3), stacked (3), december (3), mining (3), advances (3), 1145 (3), discovery (3), robust (3), special (3), nonlinear (3), baldi (3), computational (3), 2008 (3), fashion (3), mnist (3), bayes (3), 1988 (3), association (3), singular (3), elman (3), framework (3), without (3), autoassociative (3), neucom (3), neurocomputing (3), will (3), pdf (3), cheng (3), further (3), optimize (3), output (3), encoded (3), results (3), were (3), trained (3), application (3), setting (3), labels (3), exist (3), larger (3), good (3), anomalous (3), perform (3), points (3), normal (3), train (3), could (3), after (3), original (3), sigma (3), indeed (3), vectors (3), similar (3), activations (3), weights (3), span (3), dimensions (3), called (3), identity (3), pair (3), restricted (3), technique (3), objective (3), depends (3), over (3), theory (3), limit (3), daes (3), make (3), perturbations (3), caes (3), maps (3), desired (3), matrix (3), process (3), problem (3), version (3), schema (3), simply (3), define (3), sae (3), close (3), article (3), ideal (3), generate (3), just (3), parametrized (3), field (3), random (3), tools (3), languages (2), safety (2), contact (2), about (2), privacy (2), under (2), last (2), september (2), 2026 (2), categories (2), cs1 (2), maint (2), periodical (2), articles (2), total (2), pass (2), thermal (2), analyzer (2), topics (2), contrast (2), quantization (2), modulation (2), per (2), carrier (2), statistical (2), phase (2), temperature (2), channel (2), white (2), additive (2), masking (2), government (2), regulation (2), control (2), physics (2), impact (2), visual (2), social (2), act (2), mamba (2), rnn (2), perceptron (2), mlp (2), gru (2), memory (2), lstm (2), differentiable (2), turing (2), quoc (2), oriol (2), vinyals (2), ilya (2), sutskever (2), krizhevsky (2), fei (2), andrew (2), jürgen (2), schmidhuber (2), paul (2), rule (2), classifiers (2), ibm (2), large (2), generation (2), dall (2), alexnet (2), protocol (2), neuro (2), open (2), source (2), coding (2), word (2), rlhf (2), sarsa (2), sigmoid (2), tradeoff (2), projects (2), glossary (2), review (2), chinese (2), deeper (2), sequence (2), 1409 (2), kyunghyun (2), properties (2), approaches (2), predicting (2), posts (2), 5090 (2), cscita (2), 2nd (2), computing (2), designed (2), 1040 (2), 1038 (2), 2020 (2), hdl (2), biomedical (2), informatics (2), alzheimer (2), disease (2), liu (2), transactions (2), breast (2), cancer (2), song (2), various (2), isbi (2), zeng (2), wang (2), super (2), resolution (2), cybernetics (2), icdmw (2), new (2), improves (2), corrupted (2), lossy (2), score (2), know (2), 105 (2), zhou (2), 4503 (2), lecture (2), australia (2), semi (2), diagnosis (2), macao (2), plaut (2), subspaces (2), 1804 (2), 10253 (2), pierre (2), ontology (2), biology (2), graphical (2), hashing (2), adaptive (2), boosting (2), torch (2), diederik (2), kingma (2), welling (2), max (2), 1312 (2), 5947 (2), scholarpedia (2), 5786 (2), 507 (2), 1658773 (2), 16873662 (2), 1126 (2), 1127647 (2), 2006sci (2), 504h (2), 504 (2), perceptrons (2), cottrell (2), munro (2), society (2), internal (2), 0001 (2), harrison (2), continuous (2), cambridge (2), 1986 (2), parallel (2), distributed (2), 262 (2), hornik (2), 1989 (2), 0893 (2), 6080 (2), oja (2), 1982 (2), 273 (2), neuron (2), free (2), icml (2), chemical (2), 0098 (2), 1354 (2), link (2), cite (2), our (2), 1561 (2), trends (2), mark (2), 1991 (2), yuan (2), wei (2), words (2), media (2), springer (2), criterion (2), research (2), courville (2), aaron (2), bank (2), dor (2), koenigstein (2), giryes (2), raja (2), 2023 (2), 24627 (2), 24628 (2), 9_16 (2), handbook (2), references (2), reading (2), see (2), because (2), help (2), errors (2), inherent (2), difficulty (2), accurately (2), behavior (2), real (2), referred (2), sequences (2), procedure (2), generated (2), specific (2), still (2), online (2), strategies (2), experimentally (2), drug (2), experiments (2), between (2), characteristics (2), best (2), considered (2), rare (2), events (2), very (2), these (2), recent (2), certain (2), autoencoding (2), reconstructing (2), understood (2), those (2), when (2), should (2), cases (2), reconstruct (2), validation (2), since (2), become (2), spaces (2), mapping (2), support (2), query (2), less (2), powerful (2), lower (2), solution (2), orthogonal (2), projection (2), equal (2), multi (2), smaller (2), both (2), identical (2), cost (2), early (2), discrete (2), like (2), etc (2), developed (2), module (2), generalized (2), showed (2), layered (2), exponentially (2), number (2), through (2), subset (2), minimize (2), infinitesimal (2), resist (2), encourage (2), obtained (2), aim (2), frobenius (2), jacobian (2), want (2), itself (2), expected (2), thus (2), neighborhood (2), even (2), fraction (2), chosen (2), include (2), type (2), noisy (2), circ (2), inputs (2), giving (2), cdot (2), upon (2), let (2), need (2), measure (2), times (2), infty (2), instead (2), following (2), top (2), yellow (2), variants (2), codes (2), fewer (2), active (2), vaes (2), families (2), different (2), receives (2), capacity (2), overcomplete (2), undercomplete (2), really (2), appears (2), parts (2), reconstructs (2), interpretation (2), euclidean (2), arg (2), write (2), refer (2), decoded (2), family (2), rightarrow (2), definition (2), subsequent (2), problems (2), curve (2), net (2), forest (2), factor (2), structured (2), appearance (2), upload (2), file (2), changes (2), links (2), read (2), create (2), account (2), donate (2), menu (2), topic, mobile, cookie, statement, statistics, developers, conduct, legal, contacts, disclaimers, available, additional, apply, site, you, agree, registered, trademark, profit, organization, wikimedia, foundation, creative, commons, attribution, sharealike, license, rendered, parsoid, edited, utc, dmy, dates, matches, retrieved, https, org, index, php, title, oldid, 1375973936, prior, shrinkage, fields, bm3d, block, matching, filtering, bilateral, anisotropic, blur, wavelet, median, denoise, radiation, spectrum, generator, colors, acoustics, cnr, sqnr, sinr, plus, snr, sinad, mer, symbol, bit, dbrnc, receiver, ratios, pseudorandom, nvh, vibration, harshness, spectral, shaping, floor, figure, impulse, pulse, resistance, circuit, coherent, worley, pink, johnson, nyquist, jitter, infrasound, grey, flicker, cosmic, burst, brownian, background, atmospheric, awgn, class, transportation, sound, ships, rooms, radio, environment, electronics, buildings, power, measurement, cancellation, acoustic, quieting, telecommunications, category, workplace, warfare, military, art, games, marketing, chatbot, psychosis, healthcare, fiction, education, engine, explainable, environmental, competition, arms, race, anthropomorphism, winter, slop, literacy, infrastructure, effect, center, bubble, boom, economic, opposition, centers, propaganda, virtual, politician, precautionary, nationalism, ethics, elections, takeover, alignment, cold, war, political, graph, gnn, adversarial, highway, residual, cnn, echo, state, gated, unit, long, term, vit, françois, chollet, daniel, kokotajlo, jan, leike, mustafa, suleyman, schulman, aidan, gomez, shazeer, ashish, vaswani, andrej, karpathy, silver, demis, hassabis, goodnight, graves, stephen, grossberg, lotfi, zadeh, yann, lecun, hopfield, werbos, seppo, linnainmaa, seymour, papert, joseph, weizenbaum, bernard, widrow, frank, rosenblatt, oliver, selfridge, herbert, simon, cliff, shaw, allen, newell, nathaniel, rochester, mccarthy, marvin, minsky, takeo, kanade, kunihiko, fukushima, shun, ichi, amari, claude, shannon, christopher, manning, von, neumann, walter, pitts, warren, sturgis, mcculloch, alan, people, yago, dbpedia, conceptnet, bases, opencog, lida, clarion, soar, reasoners, procedural, logic, programs, inference, engines, expert, deductive, robot, autogpt, action, muzero, driving, car, openai, five, alphazero, alphago, decisional, watsonx, watson, project, debater, oasis, genie, udio, suno, riffusion, music, veo, seedance, sora, kling, hailuo, runway, gen, dream, recraft, midjourney, imagen, ideogram, gpt, flux, firefly, aurora, alphafold, whisper, elevenlabs, ocr, hwr, wavenet, implementations, physical, agent2agent, hypothetical, superintelligence, asi, agi, weak, lethal, autonomous, weapons, laws, humanity, exam, companion, intelligent, game, playing, theorem, proving, actor, critic, situated, sovereign, blended, vibe, embedding, hallucination, recursive, improvement, reflection, llm, post, uncanny, valley, rag, adversary, autoregression, imitation, prompt, augmentation, initialization, gating, rectifier, softmax, batchnorm, normalization, convolution, attention, conjugate, quasi, newton, sgd, overfitting, double, hyperparameter, parameter, constraint, satisfaction, planning, concepts, lists, software, proprietary, institutions, companies, timeline, alnaseri, omar, alzubaidi, laith, himeur, yassine, timmermann, jens, 2024, design, next, eess, 2412, 13843, han, lifeng, kuang, shaohui, incorporating, radicals, character, 1805, 01565, 3215, bart, van, merrienboer, bahdanau, dzmitry, 1259, shaunak, maity, abhishek, goel, vritti, shitole, sanjay, bhattacharya, avik, instagram, lifestyle, magazine, 177, 35350962, 4381, 8066548, 174, gregory, barber, wired, molecule, exhibits, druglike, qualities, zhavoronkov, enables, rapid, identification, potent, ddr1, kinase, inhibitors, 201716327, 31477924, s41587, 019, 0224, nature, biotechnology, martinez, murcia, francisco, ortiz, andres, gorriz, juan, ramirez, javier, castillo, barnes, diego, 195187846, 31217131, 10630, 28806, jbhi, 2914970, 2020ijbhi, 17m, studying, manifold, xiang, lei, qingshan, gilmore, hannah, jianzhong, tang, jinghai, madabhushi, anant, 130, 26208307, 4729702, pmc, tmi, 2458702, 2016itmi, 119x, 119, ssae, nuclei, histopathology, tzu, hsi, sanchez, victor, hesham, eidaly, nasir, rajpoot, hybrid, curvature, types, cells, bone, marrow, trephine, biopsy, 1043, 7433130, 1172, 7950694, 14th, symposium, kun, ruxin, cuihua, tao, dacheng, coupled, 20787612, 26625442, 2168, 2267, tcyb, 2501373, 2017itcyb, 27z, gondara, lovedeep, barcelona, spain, 246, 14354973, 9781509059102, 0041, 2016arxiv160804667g, 1608, 04667, 241, 16th, workshops, buades, coll, morel, 2005, 530, 218466166, 1137, 040616024, 490, multiscale, simulation, 1301, 3468, february, sparsification, highly, 432, 440, balle, laparra, simoncelli, april, optimized, 1611, 01704, theis, lucas, shi, wenzhe, cunningham, huszár, ferenc, compressive, 1703, 00395, xiao, zhisheng, yan, qing, amit, yali, 2003, 02977, likelihood, regret, out, nalisnick, eric, matsukawa, akihiro, teh, yee, whye, gorur, dilan, lakshminarayanan, balaji, don, 1810, 09136, ribeiro, manassés, lazzaretti, andré, eugênio, lopes, heitor, silvério, videos, patrec, 016, 2018parel, 13r, pattern, letters, chong, paffenroth, randy, 674, 207557733, 4887, 3097983, 3098052, 665, 23rd, sigkdd, sakurada, mayu, yairi, takehisa, gold, coast, qld, 14613395, 3159, 2689746, 2689747, mlsda, workshop, sensory, morales, forero, bassetto, methodology, 1037, 211027131, 7281, 3804, ieem44572, 8978509, 1031, industrial, management, ieem, chicco, davide, sadowski, peter, gene, annotation, predictions, 533, 207217210, 9781450328944, 11311, 964622, 2649387, 2649442, 5th, bioinformatics, bcb, ruslan, section, 0888, 613x, ijar, 006, 969, approximate, github, schwenk, holger, 1997, ackley, sejnowski, march, 1985, 169, s0364, 0213, 80012, 147, generating, faces, boesen, larsen, sonderby, blog, html, 6114, overview, 117, 11715509, 25462637, neunet, 003, 1404, 7828, 4249, 2009schpj, 5947h, bourlard, kamp, 294, 206775335, 3196773, bf00332918, 291, biological, garrison, annual, meeting, gray, scale, extensional, programming, jeffrey, 1626, 3372872, 4966, 1121, 395916, 1988asaj, 1615e, 1615, acoustical, america, connectionist, university, dissertation, rumelhart, mcclelland, 29140, 7551, mitpress, 5236, 001, explorations, microstructure, cognition, kurt, 90014, minima, erkki, 7153672, 1432, 1416, bf00275687, 267, simplified, aistats, 448, 455, yingbo, arpit, devansh, nwogu, ifeoma, govindaraju, venu, 1405, 1380, elad, abid, abubakar, balin, muhammad, fatih, zou, 1901, 09346, zemel, richard, 1993, morgan, kaufmann, helmholtz, rifai, explicit, invariance, extraction, 1992, neutral, 328, 80051, computers, nianyin, zhang, hong, baoye, weibo, yurong, dobaie, abdullah, expression, via, 649, 0925, 2312, 043, 643, nair, vinod, nips, usa, curran, associates, 1347, 9781615679119, 1339, 22nd, object, nets, cs294a, notes, makhzani, alireza, frey, brendan, 5663, books, brain, 046506192, master, quest, ultimate, remake, domingos, pedro, 207178999, 23946944, 2200000006, 1795, 243, 1002, aic, 690370209, 1991aiche, 233k, 233, aiche, chen, jiun, daw, ran, 055, 139, huang, jau, chi, yang, wen, chie, 3150, 030, perception, géron, aurélien, canada, reilly, 740, 739, hands, scikit, keras, tensorflow, berlin, heidelberg, transforming, introduction, 392, 174802445, 2200000056, 2019arxiv190602691k, 1906, 02691, 307, vincent, pascal, larochelle, hugo, 2010, 3408, 3371, 0262035613, rokach, lior, maimon, oded, shmueli, erez, eds, 374, 353, computation, mass, 03561, cham, publishing, dictionary, important, resilient, impairments, crucial, transmitting, minimizing, addition, solve, several, limitations, designing, complex, channels, unlike, does, match, texts, treated, side, target, incorporate, rarely, done, due, availability, linguistic, recently, produced, promising, helpful, advertising, molecules, validated, mice, demanding, contexts, well, assisted, modelling, relation, decline, mri, preprocessing, outperformed, proved, competitive, against, jpeg, 2000, analyze, flagged, true, sense, metrics, fundamental, challenge, comes, gathered, imbalanced, indicating, introducing, estimates, confidence, intervals, evaluation, literature, shown, counterintuitively, consequently, able, reliably, intuitively, considering, rein, reconstructions, far, away, region, lie, axis, replicate, salient, constraints, described, previously, encouraged, precisely, reproduce, frequently, observed, facing, worsen, instances, others, frequency, observation, contribution, ignored, failing, unfamiliar, detect, recorded, percentile, taken, flag, estimate, correctly, asymptotically, grows, extreme, potentially, big, uncertainty, choice, estimated, implies, benefits, particularly, kinds, proposed, 2007, produce, database, stored, returning, slightly, flipping, bits, hash, potential, resides, linearity, allowing, generalizations, significantly, strongly, spanned, onto, generally, yet, recovered, them, 28x28pixel, come, improve, hallmark, place, semantically, near, pretrained, stack, initialize, gradually, until, hitting, bottleneck, neurons, resulting, yielded, qualitatively, easier, interpret, clearly, separating, clusters, rbms, plot, being, apart, rotation, selects, orientation, reflections, invariant, rotations, modern, associative, days, terminology, uncertain, associating, diabolo, date, 1990s, concept, became, widely, 2010s, involved, modules, generators, ais, immediately, resurgence, 1980s, suggested, put, mode, implemented, pairs, noted, subsequently, termed, researchers, debated, whether, whole, together, global, along, representative, layerwise, success, heavily, adopted, his, involves, treating, neighboring, pretraining, approximates, fine, tune, allows, show, whose, eigenvectors, yield, shallow, decrease, amount, needed, reduce, representing, often, decoders, offers, schematic, fully, connected, forces, consist, user, specified, uses, allow, gradients, selector, makes, categorical, relaxation, seeks, includes, expressed, represents, compressed, denotes, advanced, leverages, specifically, posits, provides, shortest, combined, ensure, compact, interpretable, finite, sized, extracted, additionally, stochastically, having, versions, point, analytically, penalizing, evaluated, samples, adds, ness, square, respect, understand, note, fact, property, leads, perhaps, pictures, look, exactly
Text of the page (random words):
tzmann machine restricted gan diffusion model som convolutional neural network u net lenet alexnet deepdream neural field neural radiance field physics informed neural networks transformer vision mamba spiking neural network memtransistor electrochemical ram ecram reinforcement learning q learning policy gradient sarsa temporal difference td multi agent self play learning with humans active learning crowdsourcing human in the loop mechanistic interpretability rlhf model diagnostics coefficient of determination confusion matrix learning curve roc curve mathematical foundations kernel machines bias variance tradeoff computational learning theory empirical risk minimization occam learning pac learning statistical learning vc theory topological deep learning journals and conferences aaai cvpr eccv ecml pkdd emnlp iccv neurips icml iclr ijcai ml jmlr related articles glossary of artificial intelligence list of datasets for machine learning research list of datasets in computer vision and image processing outline of machine learning v t e an autoencoder is a type of artificial neural network used to learn efficient codings of unlabeled data unsupervised learning an autoencoder learns two functions an encoding function that transforms the input data and a decoding function that recreates the input data from the encoded representation the autoencoder learns an efficient representation encoding for a set of data typically for dimensionality reduction to generate lower dimensional embeddings for subsequent use by other machine learning algorithms 1 variants exist which aim to make the learned representations assume useful properties 2 examples are regularized autoencoders sparse denoising and contractive autoencoders which are effective in learning representations for subsequent classification tasks 3 and variational autoencoders which can be used as generative models 4 autoencoders are applied to many problems including facial recognition 5 feature detection 6 anomaly detection and learning the meaning of words 7 8 in terms of data synthesis autoencoders can also be used to randomly generate new data that is similar to the input training data 6 mathematical principles edit definition edit an autoencoder is defined by the following components two sets the space of encoded messages z displaystyle mathcal z the space of decoded messages x displaystyle mathcal x typically x displaystyle mathcal x and z displaystyle mathcal z are euclidean spaces that is x r m z r n displaystyle mathcal x mathbb r m mathcal z mathbb r n with m n displaystyle m n two parametrized families of functions the encoder family e ϕ x z displaystyle e_ phi mathcal x rightarrow mathcal z parametrized by ϕ displaystyle phi the decoder family d θ z x displaystyle d_ theta mathcal z rightarrow mathcal x parametrized by θ displaystyle theta for any x x displaystyle x in mathcal x we usually write z e ϕ x displaystyle z e_ phi x and refer to it as the code the latent variable latent representation latent vector etc conversely for any z z displaystyle z in mathcal z we usually write x d θ z displaystyle x d_ theta z and refer to it as the decoded message usually both the encoder and the decoder are defined as multilayer perceptrons mlps for example a one layer mlp encoder e ϕ displaystyle e_ phi is e ϕ x σ w x b displaystyle e_ phi mathbf x sigma wx b where σ displaystyle sigma is an element wise activation function w displaystyle w is a weight matrix and b displaystyle b is a bias vector training an autoencoder edit an autoencoder by itself is simply a tuple of two functions to judge its quality we need a task a task is defined by a reference probability distribution μ r e f displaystyle mu _ ref over x displaystyle mathcal x and a reconstruction quality function d x x 0 displaystyle d mathcal x times mathcal x to 0 infty such that d x x displaystyle d x x measures how much x displaystyle x differs from x displaystyle x with those we can define the loss function for the autoencoder as l θ ϕ e x μ r e f d x d θ e ϕ x displaystyle l theta phi mathbb mathbb e _ x sim mu _ ref d x d_ theta e_ phi x the optimal autoencoder for the given task μ r e f d displaystyle mu _ ref d is then arg min θ ϕ l θ ϕ displaystyle arg min _ theta phi l theta phi the search for the optimal autoencoder can be accomplished by any mathematical optimization technique but usually by gradient descent this search process is referred to as training the autoencoder in most situations the reference distribution is just the empirical distribution given by a dataset x 1 x n x displaystyle x_ 1 x_ n subset mathcal x so that μ r e f 1 n i 1 n δ x i displaystyle mu _ ref frac 1 n sum _ i 1 n delta _ x_ i where δ x i displaystyle delta _ x_ i is the dirac measure the quality function is just l 2 displaystyle l 2 loss d x x x x 2 2 displaystyle d x x x x _ 2 2 and 2 displaystyle cdot _ 2 is the euclidean norm then the problem of searching for the optimal autoencoder is just a least squares optimization min θ ϕ l θ ϕ where l θ ϕ 1 n i 1 n x i d θ e ϕ x i 2 2 displaystyle min _ theta phi l theta phi qquad text where l theta phi frac 1 n sum _ i 1 n x_ i d_ theta e_ phi x_ i _ 2 2 interpretation edit an autoencoder has two main parts an encoder that maps the message to a code and a decoder that reconstructs the message from the code an optimal autoencoder would perform as close to perfect reconstruction as possible with close to perfect defined by the reconstruction quality function d displaystyle d the simplest way to perform the copying task perfectly would be to duplicate the signal to suppress this behavior the code space z displaystyle mathcal z usually has fewer dimensions than the message space x displaystyle mathcal x such an autoencoder is called undercomplete it can be interpreted as compressing the message or reducing its dimensionality 9 10 at the limit of an ideal undercomplete autoencoder every possible code z displaystyle z in the code space is used to encode a message x displaystyle x that really appears in the distribution μ r e f displaystyle mu _ ref and the decoder is also perfect d θ e ϕ x x displaystyle d_ theta e_ phi x x this ideal autoencoder can then be used to generate messages indistinguishable from real messages by feeding its decoder arbitrary code z displaystyle z and obtaining d θ z displaystyle d_ theta z which is a message that really appears in the distribution μ r e f displaystyle mu _ ref if the code space z displaystyle mathcal z has dimension larger than overcomplete or equal to the message space x displaystyle mathcal x or the hidden units are given enough capacity an autoencoder can learn the identity function and become useless however experimental results found that overcomplete autoencoders might still learn useful features 11 in the ideal setting the code dimension and the model capacity could be set on the basis of the complexity of the data distribution to be modeled a standard way to do so is to add modifications to the basic autoencoder to be detailed below 2 variations edit variational autoencoder vae edit the basic scheme of a variational autoencoder the model receives x displaystyle x as input the encoder compresses it into the latent space the decoder receives as input the information sampled from the latent space and produces x displaystyle x as similar as possible to x displaystyle x main article variational autoencoder variational autoencoders vaes belong to the families of variational bayesian methods despite the architectural similarities with basic autoencoders vaes are architected with different goals and have a different mathematical formulation the latent space is in this case composed of a mixture of distributions instead of fixed vectors given an input dataset x displaystyle x characterized by an unknown probability function p x displaystyle p x and a multivariate latent encoding vector z displaystyle z the objective is to model the data as a distribution p θ x displaystyle p_ theta x with θ displaystyle theta defined as the set of the network parameters so that p θ x z p θ x z d z displaystyle p_ theta x int _ z p_ theta x z dz sparse autoencoder sae edit inspired by the sparse coding hypothesis in neuroscience sparse autoencoders sae are variants of autoencoders such that the codes e ϕ x displaystyle e_ phi x for messages tend to be sparse codes that is e ϕ x displaystyle e_ phi x is close to zero in most entries sparse autoencoders may include more rather than fewer hidden units than inputs but only a small number of the hidden units are allowed to be active at the same time 12 encouraging sparsity improves performance on classification tasks 13 simple schema of a single layer sparse autoencoder the hidden nodes in bright yellow are activated while the light yellow ones are inactive the activation depends on the input there are two main ways to enforce sparsity one way is to simply clamp all but the highest k activations of the latent code to zero this is the k sparse autoencoder 13 the k sparse autoencoder inserts the following k sparse function in the latent layer of a standard autoencoder f k x 1 x n x 1 b 1 x n b n displaystyle f_ k x_ 1 x_ n x_ 1 b_ 1 x_ n b_ n where b i 1 displaystyle b_ i 1 if x i displaystyle x_ i ranks in the top k and 0 otherwise backpropagating through f k displaystyle f_ k is simple set gradient to 0 for b i 0 displaystyle b_ i 0 entries and keep gradient for b i 1 displaystyle b_ i 1 entries this is essentially a generalized relu function 13 the other way is a relaxed version of the k sparse autoencoder instead of forcing sparsity we add a sparsity regularization loss then optimize for min θ ϕ l θ ϕ λ l sparse θ ϕ displaystyle min _ theta phi l theta phi lambda l_ text sparse theta phi where λ 0 displaystyle lambda 0 measures how much sparsity we want to enforce 14 let the autoencoder architecture have k displaystyle k layers to define a sparsity regularization loss we need a desired sparsity ρ k displaystyle hat rho _ k for each layer a weight w k displaystyle w_ k for how much to enforce each sparsity and a function s 0 1 0 1 0 displaystyle s 0 1 times 0 1 to 0 infty to measure how much two sparsities differ for each input x displaystyle x let the actual sparsity of activation in each layer k displaystyle k be ρ k x 1 n i 1 n a k i x displaystyle rho _ k x frac 1 n sum _ i 1 n a_ k i x where a k i x displaystyle a_ k i x is the activation in the i displaystyle i th neuron of the k displaystyle k th layer upon input x displaystyle x the sparsity loss upon input x displaystyle x for one layer is s ρ k ρ k x displaystyle s hat rho _ k rho _ k x and the sparsity regularization loss for the entire autoencoder is the expected weighted sum of sparsity losses l sparse θ ϕ e x μ x k 1 k w k s ρ k ρ k x displaystyle l_ text sparse theta phi mathbb mathbb e _ x sim mu _ x left sum _ k in 1 k w_ k s hat rho _ k rho _ k x right typically the function s displaystyle s is either the kullback leibler kl divergence as 13 14 15 16 s ρ ρ k l ρ ρ ρ log ρ ρ 1 ρ log 1 ρ 1 ρ displaystyle s rho hat rho kl rho hat rho rho log frac rho hat rho 1 rho log frac 1 rho 1 hat rho or the l1 loss as s ρ ρ ρ ρ displaystyle s rho hat rho rho hat rho or the l2 loss as s ρ ρ ρ ρ 2 displaystyle s rho hat rho rho hat rho 2 alternatively the sparsity regularization loss may be defined without reference to any desired sparsity but simply force as much sparsity as possible in this case one can define the sparsity regularization loss as l sparse θ ϕ e x μ x k 1 k w k h k displaystyle l_ text sparse theta phi mathbb mathbb e _ x sim mu _ x left sum _ k in 1 k w_ k h_ k right where h k displaystyle h_ k is the activation vector in the k displaystyle k th layer of the autoencoder the norm displaystyle cdot is usually the l1 norm giving the l1 sparse autoencoder or the l2 norm giving the l2 sparse autoencoder denoising autoencoder dae edit a schema of a denoising autoencoder denoising autoencoders dae try to achieve a good representation by changing the reconstruction criterion 2 3 a dae originally called a robust autoassociative network by mark a kramer 17 is trained by intentionally corrupting the inputs of a standard autoencoder during training a noise process is defined by a probability distribution μ t displaystyle mu _ t over functions t x x displaystyle t mathcal x to mathcal x that is the function t displaystyle t takes a message x x displaystyle x in mathcal x and corrupts it to a noisy version t x displaystyle t x the function t displaystyle t is selected randomly with a probability distribution μ t displaystyle mu _ t given a task μ ref d displaystyle mu _ text ref d the problem of training a dae is the optimization problem min θ ϕ l θ ϕ e x μ x t μ t d x d θ e ϕ t x displaystyle min _ theta phi l theta phi mathbb mathbb e _ x sim mu _ x t sim mu _ t d x d_ theta circ e_ phi circ t x that is the optimal dae should take any noisy message and attempt to recover the original message without noise thus the name denoising usually the noise process t displaystyle t is applied only during training and testing not during downstream use the use of dae depends on two assumptions there exist representations to the messages that are relatively stable and robust to the type of noise we are likely to encounter the said representations capture structures in the input distribution that are useful for our purposes 3 example noise processes include additive isotropic gaussian noise masking noise a fraction of the input is randomly chosen and set to 0 salt and pepper noise a fraction of the input is randomly chosen and randomly set to its minimum or maximum value 3 contractive autoencoder cae edit a contractive autoencoder cae adds the contractive regularization loss to the standard autoencoder loss min θ ϕ l θ ϕ λ l cont θ ϕ displaystyle min _ theta phi l theta phi lambda l_ text cont theta phi where λ 0 displaystyle lambda 0 measures how much contractive ness we want to enforce the contractive regularization loss itself is defined as the expected square of frobenius norm of the jacobian matrix of the encoder activations with respect to the input l cont θ ϕ e x μ r e f x e ϕ x f 2 displaystyle l_ text cont theta phi mathbb e _ x sim mu _ ref nabla _ x e_ phi x _ f 2 to understand what l cont displaystyle l_ text cont measures note the fact e ϕ x δ x e ϕ x 2 x e ϕ x f δ x 2 displaystyle e_ phi x delta x e_ phi x _ 2 leq nabla _ x e_ phi x _ f delta x _ 2 for any message x x displaystyle x in mathcal x and small variation δ x displaystyle delta x in it thus if x e ϕ x f 2 displaystyle nabla _ x e_ phi x _ f 2 is small it means that a small neighborhood of the message maps to a small neighborhood of its code this is a desired property as it means small variation in the message leads to small perhaps even zero variation in its code like how two pictures may look the same even if they are not exactly the same the cae can be understood as an infinitesimal limit o...
|