Meta tags:
Headings (most frequently used words):
methods, predicting, labels, algorithms, of, categorical, for, learning, real, valued, pattern, recognition, sequence, labeling, sequences, contents, overview, problem, statement, uses, see, also, references, further, reading, external, links, probabilistic, classifiers, number, important, feature, variables, frequentist, or, bayesian, approach, to, classification, clustering, classifying, and, ensemble, supervised, meta, combining, multiple, together, general, arbitrarily, structured, sets, multilinear, subspace, multidimensional, data, using, tensor, representations, regression,
Text of the page (most frequently used words):
the (199), and (90), #pattern (73), boldsymbol (68), #recognition (67), learning (66), for (51), displaystyle (47), data (43), theta (42), algorithms (31), that (31), with (26), machine (26), from (25), labels (24), edit (23), feature (23), this (22), label (21), analysis (20), classification (20), are (20), methods (19), neural (19), output (19), can (17), regression (16), which (16), archived (14), predicting (14), input (14), mathcal (14), may (13), all (13), statistical (13), article (13), valued (13), probability (13), probabilistic (13), given (13), doi (12), used (12), possible (12), value (12), list (11), other (11), computer (11), matching (11), vector (11), into (11), real (11), bayesian (11), mathbf (11), function (11), problem (11), 2019 (10), theory (10), journal (10), 978 (10), isbn (10), image (10), systems (10), processing (10), model (10), categorical (10), supervised (10), training (10), set (10), features (10), using (9), retrieved (9), automatic (9), information (9), original (9), networks (9), sequence (9), labeling (9), main (9), clustering (9), has (9), some (9), approach (9), wikipedia (8), non (8), general (8), new (8), example (8), models (8), number (8), also (8), unsupervised (8), classifier (8), one (8), values (8), instance (8), text (7), articles (7), template (7), wayback (7), different (7), cite (7), class (7), known (7), such (7), network (7), two (7), vectors (7), linear (7), selection (7), type (7), not (7), algorithm (7), between (7), then (7), each (7), patterns (7), statistics (6), use (6), was (6), link (6), links (6), uses (6), wiley (6), autonomous (6), 2018 (6), vision (6), discriminant (6), large (6), where (6), bayes (6), many (6), but (6), procedure (6), spam (6), left (6), instances (6), loss (6), labeled (6), toggle (5), search (5), terms (5), page (5), cs1 (5), parameters (5), short (5), multiple (5), international (5), help (5), s2cid (5), 1109 (5), bibcode (5), ieee (5), transactions (5), 2nd (5), issn (5), deep (5), techniques (5), error (5), mining (5), prior (5), process (5), only (5), its (5), random (5), logistic (5), particular (5), detection (5), based (5), extraction (5), frequentist (5), defined (5), right (5), estimate (5), typically (5), best (5), well (5), more (5), when (5), than (5), assigns (5), been (5), confidence (5), contents (4), view (4), statement (4), about (4), categories (4), pages (4), cleanup (4), needing (4), references (4), computational (4), fields (4), technology (4), learn (4), computing (4), artificial (4), research (4), york (4), classifiers (4), parameter (4), introduction (4), speech (4), further (4), shape (4), vehicle (4), verification (4), part (4), applications (4), 2020 (4), 2007 (4), digital (4), signal (4), medical (4), form (4), common (4), simple (4), mathematical (4), knowledge (4), unknown (4), perception (4), diagnosis (4), term (4), inputs (4), see (4), items (4), maximum (4), markov (4), sequences (4), pca (4), principal (4), component (4), structured (4), ensemble (4), meta (4), together (4), kernel (4), classifying (4), estimation (4), decision (4), being (4), same (4), their (4), examples (4), any (4), email (4), priori (4), probabilities (4), estimated (4), int (4), frac (4), prod (4), instead (4), over (4), find (4), correct (4), hand (4), simply (4), rightarrow (4), dimensionality (4), consisting (4), field (4), engineering (4), hide (4), move (4), sidebar (4), policy (3), available (3), additional (3), september (3), hidden (3), deprecated (3), 2014 (3), unsourced (3), maint (3), names (3), authors (3), description (3), control (3), programming (3), software (3), applied (3), intelligence (3), 2008 (3), tutorial (3), 471 (3), prediction (3), approaches (3), distributions (3), gaussian (3), psychology (3), 1016 (3), 2017 (3), automated (3), how (3), single (3), speaker (3), face (3), 2006 (3), images (3), mean (3), matrix (3), pdf (3), make (3), datasets (3), branch (3), sets (3), them (3), lists (3), multilinear (3), multidimensional (3), hierarchical (3), mixture (3), means (3), machines (3), multi (3), note (3), name (3), whether (3), generative (3), closely (3), related (3), first (3), identification (3), observation (3), max (3), according (3), distinction (3), what (3), empirical (3), observations (3), computed (3), collected (3), these (3), dataset (3), rule (3), operatorname (3), perform (3), rate (3), factor (3), sort (3), case (3), often (3), assigning (3), sometimes (3), they (3), small (3), have (3), word (3), integer (3), attempts (3), combination (3), regularities (3), overview (3), account (3), sentence (3), create (3), tools (3), subsection (3), languages (2), table (2), contact (2), privacy (2), wikimedia (2), commons (2), license (2), last (2), 2026 (2), errors (2), displaying (2), descriptions (2), redirect (2), targets (2), webarchive (2), statements (2), 2011 (2), wikidata (2), study (2), org (2), libraries (2), bias (2), neuromorphic (2), differentiable (2), improved (2), fast (2), open (2), association (2), external (2), springer (2), oclc (2), citeseerx (2), robert (2), 2000 (2), review (2), duda (2), richard (2), hart (2), peter (2), stork (2), david (2), 05669 (2), san (2), francisco (2), morgan (2), kaufmann (2), publishers (2), expert (2), 1990 (2), press (2), reading (2), distributional (2), per (2), cool (2), 2012 (2), revision (2), ifac (2), april (2), 6670 (2), driven (2), cars (2), arxiv (2), way (2), science (2), 0035 (2), sae (2), technical (2), papnet (2), cervical (2), screening (2), sahidullah (2), saha (2), goutam (2), march (2), iet (2), challenges (2), 2016 (2), com (2), plate (2), man (2), cybernetics (2), practice (2), optimization (2), 102795 (2), covariance (2), 1987 (2), optimum (2), tradeoff (2), 1972 (2), computers (2), technique (2), analyzing (2), events (2), sensory (2), numerical (2), structure (2), box (2), aided (2), system (2), viewed (2), please (2), time (2), recurrent (2), entropy (2), conditional (2), components (2), ica (2), independent (2), filters (2), subspace (2), tensor (2), representations (2), arbitrarily (2), experts (2), bagging (2), boosting (2), combining (2), expression (2), support (2), perceptrons (2), naive (2), trees (2), nature (2), categorized (2), discriminative (2), identify (2), objects (2), humans (2), made (2), concerns (2), produce (2), stimuli (2), there (2), etc (2), navigation (2), authentication (2), include (2), character (2), application (2), relative (2), pressure (2), were (2), content (2), cad (2), typical (2), human (2), handwriting (2), weighted (2), objective (2), posteriori (2), considered (2), does (2), posterior (2), likelihood (2), point (2), map (2), arg (2), learned (2), conflicting (2), context (2), regularization (2), mathematically (2), rather (2), directly (2), combined (2), follows (2), mapping (2), approximates (2), either (2), needs (2), specific (2), resulting (2), incorrect (2), goal (2), minimize (2), expectation (2), taken (2), distribution (2), ground (2), truth (2), amount (2), test (2), equivalent (2), zero (2), attempt (2), reduce (2), work (2), less (2), after (2), while (2), complexity (2), because (2), explored (2), medium (2), important (2), variables (2), larger (2), correspondingly (2), associated (2), described (2), formally (2), blood (2), ordinal (2), space (2), task (2), inherent (2), pre (2), classes (2), commonly (2), community (2), generally (2), assumes (2), definition (2), determine (2), unlabeled (2), semi (2), occam (2), actions (2), modern (2), capabilities (2), conference (2), curve (2), self (2), reinforcement (2), net (2), forest (2), anomaly (2), reduction (2), shot (2), sources (2), citations (2), appearance (2), upload (2), file (2), changes (2), history (2), read (2), bahasa (2), log (2), donate (2), menu (2), add, topic, mobile, cookie, developers, code, conduct, legal, safety, contacts, disclaimers, under, apply, site, you, agree, registered, trademark, profit, organization, foundation, inc, creative, attribution, sharealike, rendered, parsoid, edited, utc, via, module, annotated, missing, periodical, archiveis, january, formal, sciences, https, index, php, title, pattern_recognition, oldid, 1374822092, yale, lux, israel, czech, republic, japan, united, states, national, gnd, authority, databases, portals, mindspore, flux, jax, theano, scikit, keras, pytorch, tensorflow, spinnaker, memristor, vpu, tpu, ipu, hardware, inductive, ricci, calculus, differentiation, manifold, geometry, intended, source, platform, sharing, project, 2004, society, info, web, sites, kovalevsky, 1980, 852790446, 4612, 6033, introductory, introducing, basic, numeric, jain, anil, duin, mao, jianchang, 192934, 824819, 123, 8151, 2000itpam, interscience, kulikowski, casimir, weiss, sholom, 1991, 55860, 065, nets, godfried, toussaint, 1988, amsterdam, north, holland, publishing, company, 4832, 9672, morphology, schuermann, juergen, 1996, 13534, unified, hornegger, joachim, paulus, dietrich, 1999, 528, 15558, practical, fukunaga, keinosuke, boston, academic, 269851, assumption, regarding, assuming, 2013, level, attention, website, sinha, hadjiiski, mutib, 1993, 1st, workshop, intelligent, vehicles, hampshire, 340, 1474, s1474, 49322, 335, proceedings, volumes, requires, ray, baishakhi, jana, suman, pei, kexin, tian, yuchi, deeptest, testing, 2017arxiv170808559t, 1708, 08559, pickering, chris, engineer, paving, fully, gerdes, christian, kegelman, john, kapania, nitin, brown, matthew, spielberg, nathan, eaaw1975, 89616974, 33137751, pmid, 2470, 9476, 1126, scirobotics, aaw1975, robotics, high, performance, driving, 4271, saemobilus, development, strategy, camera, paper, mobilus, archive, today, poddar, arnab, 101, 1049, bmt, 0065, biometrics, utterances, trends, opportunities, companion, chapter, textbook, http, anpr, impedovo, donato, pirlo, giuseppe, 635, 1558, 2442, tsmcc, 923866, 609, reviews, signature, state, art, brunelli, 2009, 470, 51706, book, 2001, sarangi, susanta, filterbank, 220665533, dsp, 2020dsp, 10402795s, 10729, 104, milewski, govindaraju, venu, 1315, october, patcog, 018, 2008patre, 1308m, 1308, binarization, handwritten, carbon, copy, consists, sigma, iman, foroutan, jack, sklansky, 198, 9871395, tsmc, 4309029, 1987itsmc, 187f, 187, isabelle, guyon, clopinet, andré, elisseeff, 2003, vol, 1157, 1182, variable, edition, chow, 1557, 9654, 1721, 6177, hdl, tit, 1970, 1054406, reject, carvalko, preston, determining, golay, marking, transforms, binary, 21050445, 223519, 1430, bishop, christopher, ian, chiswell, oxford, university, 799802313, 921562, logic, utah, edu, howard, 275, 0368, 492x, 1108, 03684920710743466, kybernetes, facts, predictions, predictive, analytics, better, skills, perceptual, interpretation, neocognitron, scientific, production, limited, grey, contextual, assisted, compound, cache, language, outputs, implementation, black, neuropsychology, adaptive, resonance, removing, incorporating, clean, embedded, contain, indiscriminate, unverified, dtw, dynamic, warping, rnns, memms, hmms, crfs, extensions, kriging, particle, kalman, mpca, averaging, bootstrap, aggregating, correlation, agglomerative, divisive, cluster, gene, layer, nearest, neighbor, nonparametric, aka, despite, comes, fact, extension, multinomial, quadratic, parametric, depend, sense, explains, receive, meaningful, thought, ways, second, proportions, hypothesis, suggests, incoming, compared, templates, long, memory, match, stimulus, identified, pandemonium, letters, selfridge, 1959, suggest, broken, down, parts, capital, having, three, horizontal, lines, vertical, line, mobility, advanced, driver, assistance, defense, various, guidance, target, cancer, breast, tumors, heart, sounds, fingerprint, voice, world, optical, method, signing, captured, stylus, overlay, starting, strokes, speed, min, acceleration, uniquely, confirm, identity, banks, offered, collect, fdic, bank, fraud, did, want, inconvenience, customers, citation, needed, within, basis, describes, supports, doctor, interpretations, findings, messages, postal, envelopes, faces, forms, subtopic, deals, several, detected, facial, origin, greek, philosophy, already, later, his, before, gained, chosen, user, moreover, experience, quantified, facilitates, seamless, intermixing, subjective, dirichlet, conjugate, beta, kant, presented, developed, tradition, entails, precisely, usage, fisher, full, selecting, query, integrating, evidence, marginal, represents, propto, evaluation, theorem, finds, simultaneously, meets, smallest, simplest, essentially, combines, favors, simpler, complex, placing, denominator, involves, summation, integration, continuously, distributed, sum, parameterized, however, inverse, recognizer, stated, maps, along, assumed, represent, accurate, filtering, representation, order, rigorously, specifying, cost, producing, neither, nor, exactly, empirically, collecting, samples, consuming, limiting, depends, predicted, sufficient, corresponds, implies, optimal, minimizes, counting, fraction, wrongly, maximizing, correctly, classified, maximize, correctness, expected, dots, transform, raw, smaller, easier, encodes, redundancy, place, easily, interpretable, subset, prune, out, redundant, irrelevant, summarizes, monotonous, total, subsets, need, intractable, bound, powerset, effectively, incorporated, tasks, partially, completely, avoids, propagation, choosing, too, low, abstain, choice, grounded, meaning, compare, against, unlike, addition, fairly, advantages, inference, piece, generated, termed, constitute, characteristics, seen, defining, points, appropriate, manipulating, angle, unordered, gender, male, female, ordered, count, occurrences, measurement, grouped, require, groups, greater, discretized, nominal, dot, product, spaces, describe, corresponding, procedures, normally, involving, speak, grouping, clusters, dimensional, terminology, refer, ecology, distance, similarity, measure, generate, provided, properly, generates, meet, objectives, generalize, usually, accordance, discussed, below, cases, razor, concerned, discovery, through, take, shifted, filter, responses, cosfire, aim, provide, reasonable, answer, most, likely, taking, variation, opposed, look, exact, matches, existing, looks, textual, included, processors, editors, regular, assignment, introduced, purpose, 1936, assign, encompasses, types, member, describing, syntactic, parse, tree, parsing, tagging, trained, discover, previously, focus, stronger, connection, business, focuses, takes, acquisition, consideration, originated, popular, leading, named, kdd, extracted, similar, confused, possess, primary, distinguish, emergent, origins, due, increased, availability, abundance, power, big, graphics, compression, bioinformatics, retrieval, outline, glossary, jmlr, ijcai, iclr, icml, neurips, iccv, emnlp, ecml, pkdd, eccv, cvpr, aaai, journals, conferences, topological, pac, risk, minimization, variance, foundations, roc, confusion, coefficient, determination, diagnostics, rlhf, mechanistic, interpretability, loop, crowdsourcing, active, play, agent, temporal, difference, sarsa, gradient, ecram, electrochemical, ram, memtransistor, spiking, mamba, transformer, physics, informed, radiance, deepdream, alexnet, lenet, convolutional, som, diffusion, gan, restricted, boltzmann, reservoir, esn, gru, lstm, feedforward, autoencoder, isolation, local, outlier, ransac, graphical, sdl, sne, pgd, nmf, lda, cca, exploratory, shift, optics, dbscan, maximization, fuzzy, cure, birch, svm, relevance, rvm, perceptron, ensembles, apprenticeship, multimodal, ontology, grammar, induction, rank, semantic, rules, automl, cleaning, density, modeling, problems, quantum, neuro, symbolic, curriculum, batch, online, few, transfer, paradigms, series, remove, message, material, challenged, jstor, scholar, books, newspapers, news, removed, adding, reliable, improve, cognitive, disambiguation, free, encyclopedia, item, wikiversity, projects, printable, version, download, print, export, switch, legacy, parser, get, shortened, url, permanent, here, english, talk, tiếng, việt, українська, türkçe, tagalog, ไทย, српски, srpski, srpskohrvatski, српскохрватски, русский, português, polski, norsk, bokmål, nederlands, melayu, മലയാളം, 한국어, қазақша, jawa, 日本語, italiano, indonesia, հայերեն, hrvatski, עברית, français, suomi, فارسی, español, ελληνικά, deutsch, català, azərbaycanca, العربية, top, personal, special, recent, portal, contribute, current, jump,
Text of the page (random words):
play learning with humans active learning crowdsourcing human in the loop mechanistic interpretability rlhf model diagnostics coefficient of determination confusion matrix learning curve roc curve mathematical foundations kernel machines bias variance tradeoff computational learning theory empirical risk minimization occam learning pac learning statistical learning vc theory topological deep learning journals and conferences aaai cvpr eccv ecml pkdd emnlp iccv neurips icml iclr ijcai ml jmlr related articles glossary of artificial intelligence list of datasets for machine learning research list of datasets in computer vision and image processing outline of machine learning v t e pattern recognition is the task of assigning a class to an observation based on patterns extracted from data while similar pattern recognition pr is not to be confused with pattern machines pm which may possess pr capabilities but their primary function is to distinguish and create emergent patterns pr has applications in statistical data analysis signal processing image analysis information retrieval bioinformatics data compression computer graphics and machine learning pattern recognition has its origins in statistics and engineering some modern approaches to pattern recognition include the use of machine learning due to the increased availability of big data and a new abundance of processing power pattern recognition systems are commonly trained from labeled training data when no labeled data are available other algorithms can be used to discover previously unknown patterns kdd and data mining have a larger focus on unsupervised methods and stronger connection to business use pattern recognition focuses more on the signal and also takes acquisition and signal processing into consideration it originated in engineering and the term is popular in the context of computer vision a leading computer vision conference is named conference on computer vision and pattern recognition in machine learning pattern recognition is the assignment of a label to a given input value in statistics discriminant analysis was introduced for this same purpose in 1936 an example of pattern recognition is classification which attempts to assign each input value to one of a given set of classes for example determine whether a given email is spam pattern recognition is a more general problem that encompasses other types of output as well other examples are regression which assigns a real valued output to each input 1 sequence labeling which assigns a class to each member of a sequence of values 2 for example part of speech tagging which assigns a part of speech to each word in an input sentence and parsing which assigns a parse tree to an input sentence describing the syntactic structure of the sentence 3 pattern recognition algorithms generally aim to provide a reasonable answer for all possible inputs and to perform most likely matching of the inputs taking into account their statistical variation this is opposed to pattern matching algorithms which look for exact matches in the input with pre existing patterns a common example of a pattern matching algorithm is regular expression matching which looks for patterns of a given sort in textual data and is included in the search capabilities of many text editors and word processors overview edit further information on combination of shifted filter responses cosfire a modern definition of pattern recognition is the field of pattern recognition is concerned with the automatic discovery of regularities in data through the use of computer algorithms and with the use of these regularities to take actions such as classifying the data into different categories 4 pattern recognition is generally categorized according to the type of learning procedure used to generate the output value supervised learning assumes that a set of training data the training set has been provided consisting of a set of instances that have been properly labeled by hand with the correct output a learning procedure then generates a model that attempts to meet two sometimes conflicting objectives perform as well as possible on the training data and generalize as well as possible to new data usually this means being as simple as possible for some technical definition of simple in accordance with occam s razor discussed below unsupervised learning on the other hand assumes training data that has not been hand labeled and attempts to find inherent patterns in the data that can then be used to determine the correct output value for new data instances 5 a combination of the two that has been explored is semi supervised learning which uses a combination of labeled and unlabeled data typically a small set of labeled data combined with a large amount of unlabeled data in cases of unsupervised learning there may be no training data at all sometimes different terms are used to describe the corresponding supervised and unsupervised learning procedures for the same type of output the unsupervised equivalent of classification is normally known as clustering based on the common perception of the task as involving no training data to speak of and of grouping the input data into clusters based on some inherent similarity measure e g the distance between instances considered as vectors in a multi dimensional vector space rather than assigning each input instance into one of a set of pre defined classes in some fields the terminology is different in community ecology the term classification is used to refer to what is commonly known as clustering the piece of input data for which an output value is generated is formally termed an instance the instance is formally described by a vector of features which together constitute a description of all known characteristics of the instance these feature vectors can be seen as defining points in an appropriate multidimensional space and methods for manipulating vectors in vector spaces can be correspondingly applied to them such as computing the dot product or the angle between two vectors features typically are either categorical also known as nominal i e consisting of one of a set of unordered items such as a gender of male or female or a blood type of a b ab or o ordinal consisting of one of a set of ordered items e g large medium or small integer valued e g a count of the number of occurrences of a particular word in an email or real valued e g a measurement of blood pressure often categorical and ordinal data are grouped together and this is also the case for integer valued and real valued data many algorithms work only in terms of categorical data and require that real valued or integer valued data be discretized into groups e g less than 5 between 5 and 10 or greater than 10 probabilistic classifiers edit main article probabilistic classifier many common pattern recognition algorithms are probabilistic in nature in that they use statistical inference to find the best label for a given instance unlike other algorithms which simply output a best label often probabilistic algorithms also output a probability of the instance being described by the given label in addition many probabilistic algorithms output a list of the n best labels with associated probabilities for some value of n instead of simply a single best label when the number of possible labels is fairly small e g in the case of classification n may be set so that the probability of all possible labels is output probabilistic algorithms have many advantages over non probabilistic algorithms they output a confidence value associated with their choice note that some other algorithms may also output confidence values but in general only for probabilistic algorithms is this value mathematically grounded in probability theory non probabilistic confidence values can in general not be given any specific meaning and only used to compare against other confidence values output by the same algorithm correspondingly they can abstain when the confidence of choosing any particular output is too low 6 because of the probabilities output probabilistic pattern recognition algorithms can be more effectively incorporated into larger machine learning tasks in a way that partially or completely avoids the problem of error propagation 7 number of important feature variables edit feature selection algorithms attempt to directly prune out redundant or irrelevant features a general introduction to feature selection which summarizes approaches and challenges has been given 8 the complexity of feature selection is because of its non monotonous character an optimization problem where given a total of n displaystyle n features the powerset consisting of all 2 n 1 displaystyle 2 n 1 subsets of features need to be explored the branch and bound algorithm 9 does reduce this complexity but is intractable for medium to large values of the number of available features n displaystyle n techniques to transform the raw feature vectors feature extraction are sometimes used prior to application of the pattern matching algorithm feature extraction algorithms attempt to reduce a large dimensionality feature vector into a smaller dimensionality vector that is easier to work with and encodes less redundancy using mathematical techniques such as principal components analysis pca the distinction between feature selection and feature extraction is that the resulting features after feature extraction has taken place are of a different sort than the original features and may not easily be interpretable while the features left after feature selection are simply a subset of the original features problem statement edit the problem of pattern recognition can be stated as follows given an unknown function g x y displaystyle g mathcal x rightarrow mathcal y the ground truth that maps input instances x x displaystyle boldsymbol x in mathcal x to output labels y y displaystyle y in mathcal y along with training data d x 1 y 1 x n y n displaystyle mathbf d boldsymbol x _ 1 y_ 1 dots boldsymbol x _ n y_ n assumed to represent accurate examples of the mapping produce a function h x y displaystyle h mathcal x rightarrow mathcal y that approximates as closely as possible the correct mapping g displaystyle g for example if the problem is filtering spam then x i displaystyle boldsymbol x _ i is some representation of an email and y displaystyle y is either spam or non spam in order for this to be a well defined problem approximates as closely as possible needs to be defined rigorously in decision theory this is defined by specifying a loss function or cost function that assigns a specific value to loss resulting from producing an incorrect label the goal then is to minimize the expected loss with the expectation taken over the probability distribution of x displaystyle mathcal x in practice neither the distribution of x displaystyle mathcal x nor the ground truth function g x y displaystyle g mathcal x rightarrow mathcal y are known exactly but can be computed only empirically by collecting a large number of samples of x displaystyle mathcal x and hand labeling them using the correct value of y displaystyle mathcal y a time consuming process which is typically the limiting factor in the amount of data of this sort that can be collected the particular loss function depends on the type of label being predicted for example in the case of classification the simple zero one loss function is often sufficient this corresponds simply to assigning a loss of 1 to any incorrect labeling and implies that the optimal classifier minimizes the error rate on independent test data i e counting up the fraction of instances that the learned function h x y displaystyle h mathcal x rightarrow mathcal y labels wrongly which is equivalent to maximizing the number of correctly classified instances the goal of the learning procedure is then to minimize the error rate maximize the correctness on a typical test set for a probabilistic pattern recognizer the problem is instead to estimate the probability of each possible output label given a particular input instance i e to estimate a function of the form p l a b e l x θ f x θ displaystyle p rm label boldsymbol x boldsymbol theta f left boldsymbol x boldsymbol theta right where the feature vector input is x displaystyle boldsymbol x and the function f is typically parameterized by some parameters θ displaystyle boldsymbol theta 10 in a discriminative approach to the problem f is estimated directly in a generative approach however the inverse probability p x l a b e l displaystyle p boldsymbol x rm label is instead estimated and combined with the prior probability p l a b e l θ displaystyle p rm label boldsymbol theta using bayes rule as follows p l a b e l x θ p x l a b e l θ p l a b e l θ l all labels p x l p l θ displaystyle p rm label boldsymbol x boldsymbol theta frac p boldsymbol x rm label boldsymbol theta p rm label boldsymbol theta sum _ l in text all labels p boldsymbol x l p l boldsymbol theta when the labels are continuously distributed e g in regression analysis the denominator involves integration rather than summation p l a b e l x θ p x l a b e l θ p l a b e l θ l all labels p x l p l θ d l displaystyle p rm label boldsymbol x boldsymbol theta frac p boldsymbol x rm label boldsymbol theta p rm label boldsymbol theta int _ l in text all labels p boldsymbol x l p l boldsymbol theta operatorname d l the value of θ displaystyle boldsymbol theta is typically learned using maximum a posteriori map estimation this finds the best value that simultaneously meets two conflicting objects to perform as well as possible on the training data smallest error rate and to find the simplest possible model essentially this combines maximum likelihood estimation with a regularization procedure that favors simpler models over more complex models in a bayesian context the regularization procedure can be viewed as placing a prior probability p θ displaystyle p boldsymbol theta on different values of θ displaystyle boldsymbol theta mathematically θ arg max θ p θ d displaystyle boldsymbol theta arg max _ boldsymbol theta p boldsymbol theta mathbf d where θ displaystyle boldsymbol theta is a point estimate such as the map estimate used for θ displaystyle boldsymbol theta in point estimate evaluation and p θ d displaystyle p boldsymbol theta mathbf d the posterior probability of θ displaystyle boldsymbol theta is given by bayes theorem p θ d i 1 n p y i x i θ p θ p d i 1 n p y i x i θ p θ displaystyle p boldsymbol theta mathbf d frac left prod _ i 1 n p y_ i boldsymbol x _ i boldsymbol theta right p boldsymbol theta p mathbf d propto left prod _ i 1 n p y_ i boldsymbol x _ i boldsymbol theta right p boldsymbol theta where d x i y i i 1 n displaystyle mathbf d boldsymbol x _ i y_ i _ i 1 n represents the training dataset and p d i 1 n p y i x i θ p θ d θ displaystyle p mathbf d int left prod _ i 1 n p y_ i boldsymbol x ...
|