Meta tags:
Headings (most frequently used words):
friday, data, 2017, march, mining, in, matlab, january, 14, 2022, april, 21, wednesday, 29, 17, 18, 2016, about, me, search, links, blog, archive, implied, probabilities, new, dawn, for, local, learning, methods, analytics, summit, iii, at, harrisburg, university, of, science, and, technology, geographic, distances, quick, trip, around, the, great, circle, four, books, worth, owning,
Text of the page (most frequently used words):
the (138), and (43), odds (30), are (23), data (22), this (22), for (21), payout (21), will (17), learning (16), that (16), local (15), probability (14), with (13), which (13), house (13), implied (12), their (10), #mining (9), one (9), such (9), cosd (9), summit (9), wager (9), player (9), march (8), matlab (8), analytics (8), more (8), distance (8), lata (8), latb (8), dwinnell (7), probabilities (7), books (7), they (7), not (7), line (7), from (7), neural (7), control (7), april (6), than (6), statistics (6), statistical (6), these (6), have (6), sind (6), harrisburg (6), has (6), was (6), time (6), nearest (6), methods (6), other (6), may (5), labels (5), comments (5), posted (5), each (5), note (5), distancekilometers (5), longa (5), longb (5), new (5), university (5), event (5), been (5), some (5), networks (5), often (5), out (5), gambling (5), prediction (5), real (5), december (4), january (4), 2017 (4), 2022 (4), but (4), information (4), few (4), given (4), since (4), any (4), money (4), friday (4), trigonometry (4), longitude (4), latitude (4), great (4), circle (4), using (4), science (4), technology (4), free (4), case (4), there (4), rbf (4), hardware (4), algorithms (4), popular (4), techniques (4), well (4), components (4), model (4), make (4), favor (4), outcomes (4), statistically (4), markets (4), events (4), important (4), republican (4), party (4), senate (4), specific (4), oddsmaker (4), puts (4), theme (3), november (3), february (3), july (3), 2016 (3), links (3), fields (3), used (3), about (3), paper (3), listed (3), worth (3), largely (3), less (3), edition (3), different (3), own (3), between (3), cities (3), readers (3), might (3), situation (3), zip (3), codes (3), american (3), also (3), code (3), most (3), quick (3), second (3), isbn (3), 978 (3), round (3), 111 (3), acosd (3), random (3), reference (3), video (3), iii (3), because (3), does (3), can (3), unstructured (3), over (3), neighbors (3), neighbor (3), many (3), train (3), another (3), set (3), thing (3), training (3), all (3), among (3), them (3), would (3), predictions (3), oddsmakers (3), wagers (3), indicate (3), tend (3), same (3), total (3), 2012 (2), september (2), predictive (2), experience (2), http (2), www (2), com (2), posts (2), older (2), advisers (2), four (2), owning (2), traditional (2), machine (2), sleuth (2), rather (2), host (2), topics (2), etc (2), vary (2), consider (2), your (2), common (2), spherical (2), location (2), geographic (2), recently (2), calculate (2), interested (2), usually (2), small (2), making (2), enough (2), purposes (2), two (2), expected (2), decimal (2), degrees (2), minutes (2), formula (2), source (2), quickly (2), far (2), zeichick (2), computer (2), above (2), 9385 (2), 664274 (2), york (2), presentations (2), pennsylvania (2), food (2), hosting (2), government (2), related (2), applied (2), you (2), nice (2), additionally (2), whose (2), article (2), primer (2), https (2), icrunchdata (2), network (2), radial (2), basis (2), function (2), lazy (2), deep (2), improvement (2), while (2), interestingly (2), powerful (2), ever (2), feedforward (2), several (2), machines (2), class (2), easy (2), large (2), number (2), only (2), cases (2), need (2), comes (2), situations (2), thus (2), simply (2), slow (2), last (2), edited (2), wisdom (2), payoff (2), market (2), fractional (2), provide (2), predictors (2), houses (2), conclusion (2), example (2), fairly (2), bet (2), forces (2), driven (2), first (2), bias (2), pay (2), poor (2), betting (2), sometimes (2), loss (2), details (2), politics (2), bovada (2), after (2), democrat (2), estimated (2), average (2), half (2), leaves (2), way (2), william, awesome, inc, powered, blogger, 2006, june, october, 2007, 2008, 2009, 2010, 2013, 2014, blog, archive, mathworks, vendor, rss, feed, web, log, search, view, complete, profile, scientist, years, care, remember, worked, variety, wide, array, tools, tool, choice, find, linkedin, predictor, subscribe, atom, home, below, feel, take, perspective, opposed, exception, like, textbooks, guide, reflecting, practical, advice, respective, authors, comparatively, pages, devoted, modeling, cover, relevant, analyst, sample, size, determination, hypothesis, testing, assumptions, sampling, technique, isbns, editions, valuable, saving, skipping, latest, final, clarification, giving, blanket, endorsement, completely, disagree, ideas, see, value, use, backgrounds, perspectives, ramsey, schafer, consultant, newton, rudestam, rules, thumb, van, belle, errors, good, hardin, spatial, global, geography, wanted, locations, earth, finding, handy, solution, thought, included, postal, available, look, table, geometric, centroid, areas, identified, geographical, close, assumption, planet, perfectly, allow, calculations, precise, following, coordinates, crow, flies, technically, kilometers, seconds, conversion, needed, checked, against, verified, pairs, 1202, 3955, how, berlin, alan, published, sep, 1991, issue, language, magazine, credits, his, scientific, calculator, reverse, engineered, demystified, 2nd, stan, gibilisco, 178024, references, 118, 410825, 019394, 3940km, difference, los, angeles, 422675, 762909, 1202km, atlanta, distances, trip, around, conference, just, finished, multi, day, featuring, mix, presenters, private, sector, businesses, academia, spans, research, practice, visionary, big, picture, studies, measuring, impact, communicating, results, regrettably, unable, attend, traveling, business, held, 2015, haven, job, charge, prospect, starving, grad, student, generously, provided, recent, previous, found, bottom, apr, encourage, explore, resource, analyticssummit, harrisburgu, edu, wednesday, relentless, speed, computers, continues, technical, barriers, progress, begun, emerge, exploitation, parallelism, actually, increased, rate, acceleration, especially, mathematical, put, task, running, baroque, once, trained, days, now, affordable, desktop, fancier, fed, boosting, support, vector, forests, illustrate, trend, benefit, developments, typical, were, briefly, heyday, 1990s, much, faster, discussed, chapter, conceptually, simple, implement, demonstrates, advantages, disadvantages, handles, fraction, possible, input, approach, coordinate, complexity, having, handle, stores, future, little, fitting, done, during, gift, price, though, systems, very, execution, models, either, fire, those, spend, figuring, applies, fallen, predict, secondarily, implementation, requires, retention, author, wonders, whether, contemporary, present, opportunity, resurgence, perform, help, diversify, ensembles, users, algoprithms, analysts, looking, measure, served, investigating, solutions, easiest, scratch, directed, aha, norms, pattern, classification, dasarathy, 0818689307, 0792345848, dawn, crowd, foundation, participants, powerfully, motivated, deliver, competent, analysis, aggregating, competitive, alternative, pollsters, forecasters, similar, experts, anticipated, yet, awkward, high, importance, low, frequency, involve, special, circumstances, stymie, mechanisms, too, here, represents, membership, percent, seats, republicans, fill, exact, meaning, commercial, terms, proposed, regarding, times, conditions, neither, political, wins, split, include, stipulation, cancelled, vice, president, breaks, tie, resolution, third, casino, potential, arbitrage, keep, aligned, arguably, crowds, desire, minimize, risk, force, works, unbiased, estimates, whereas, suggested, balance, competing, despite, per, ensure, losers, winners, avoiding, pays, painfully, its, product, human, behavior, partially, emotion, gamblers, enjoy, failings, rest, humanity, spoils, aim, biases, positive, side, gain, provides, inducement, clear, thinking, long, run, tends, weed, academic, pundit, casually, talk, show, quite, when, mentioning, illustration, actual, offered, jan, midterm, election, democratic, significantly, map, together, into, space, divides, our, normally, sum, exceeds, generally, inflates, method, dealing, divide, mapping, continuum, converted, doing, arithmetic, divided, plus, continuing, hypothetical, extra, versus, believes, then, nor, work, advantage, instance, lose, dollars, nothing, putting, understand, nature, taught, courses, schedule, payment, win, typically, outcome, think, successfully, predicts, hence, upwardly, biased, rarer, overpredict, useful, mechanism, subjects, interest, beyond, sports, entertainment, current, subtleties, however, converting, called, expressed, formats, casinos, bookmakers, exploring, toolboxes,
Text of the page (random words):
in matlab exploring data mining using matlab and sometimes matlab toolboxes friday january 14 2022 implied probabilities oddsmakers provide a useful prediction mechanism for many subjects of interest beyond sports they host prediction markets for events in politics entertainment current events and other fields there are some subtleties however in converting payout odds to probabilities note that that payout odds also called payoff odds or house odds are expressed in this article as fractional odds there are several other popular formats used in casinos and on line bookmakers such as decimal odds and american odds they simply indicate the same information a different way payout odds it is important to understand the specific nature of payout odds which are not the same as odds taught in statistics courses payout odds indicate the schedule of payment for win or loss at the conclusion of a wager typically payout odds overpredict the probability of a specific outcome think about it this way the rarer the event the player successfully predicts the more the house will need to pay out hence payout odds tend to be upwardly biased if the oddsmaker believes that the probability of an event is 50 then a bet which does not favor the house nor the player would be 1 1 the house puts up 1 and the player puts up 1 to make such a situation work to the oddsmaker s advantage the payout odds might be set by the house to for instance 5 6 the house puts up only 5 while the player puts up 6 the player in this case is expected to lose an average of 0 50 dollars half the time the player leaves with 11 and the other half of the time the player leaves with nothing an average payout of 5 50 after putting up 6 implied probability payout odds can be converted to implied probabilities by doing some quick arithmetic the player s wager is divided by the total wager the house wager plus the player s wager continuing with the hypothetical payout odds of 5 6 the implied probability is 0 54 6 5 6 the extra 0 04 versus the oddsmaker s prediction of 0 5 is the statistical bias of this implied probability mapping to the probability continuum interestingly it is normally the case that the sum of the implied probabilities for real payout odds exceeds 1 0 this is because the oddsmaker generally inflates each of the payout odds to make them all in the house s favor one common method for dealing with this is to divide each of the implied probabilities by their total for an illustration using real payout odds consider an actual wager on american politics offered on jan 14 2022 at bovada https www bovada lv an on line gambling house the listed wager is which party will control the senate after the 2022 midterm election and the listed outcomes are republican with payout odds of 2 5 and democratic with payout odds of 37 20 the implied probability of republican control of the senate is 0 71 5 2 5 the implied probability of democrat control is 0 35 20 37 20 note significantly that 0 71 0 35 1 to map these implied probabilities together into the probability space one divides by their total 1 06 thus using payout odds as our model the estimated probability of republican control is 0 67 0 71 1 06 and the estimated probability of democrat control is 0 33 0 35 1 06 important details a few important details are worth mentioning first gambling is a product of human behavior betting is partially driven by emotion and gamblers enjoy the same failings as the rest of humanity this sometimes spoils their aim and statistically biases predictions of the betting markets on the positive side the gain or loss of real money provides a powerful inducement for clear thinking and in the long run tends to weed out poor predictors it s one thing for an academic or pundit to casually make a prediction in a paper or on a talk show it s quite another thing to do so when one s own money is on the line second payout odds are arguably driven by two forces 1 the wisdom of crowds and 2 the desire of the house to minimize risk the first force works in favor of statistically unbiased estimates whereas the second one may bias payout odds it has been suggested that oddsmakers will set odds to balance wagers on competing outcomes despite making less money per wager this would ensure that the losers pay the winners avoiding situations in which the house pays out painfully on its own poor predictions third payout odds may vary from one casino or on line house to another but market forces and the potential for arbitrage tend to keep them fairly well aligned last commercial gambling houses tend to be fairly specific in the terms of their proposed wagers regarding times conditions etc in the real example given above in the event that neither political party wins control there is a 50 50 split the wager will include a stipulation that the bet is cancelled or the party of the vice president breaks the tie or some other such specific resolution is used note too the exact meaning of such wagers in the real example related here the 0 67 represents the probability that the republican party will control have more than 50 of the membership of the senate it does not indicate the percent of senate seats which the republicans will fill conclusion gambling markets are at their foundation information markets participants are powerfully motivated to deliver competent analysis by aggregating such predictions oddsmakers provide a competitive alternative to more traditional predictors such as pollsters forecasters and similar experts additionally the events whose outcomes are anticipated by gambling houses are often important yet statistically awkward high importance low frequency events that often involve special one time circumstances which can stymie other prediction mechanisms posted by will dwinnell at 11 11 no comments labels fractional odds house odds implied probability information market payoff odds payout odds probability wisdom of the crowd friday april 21 2017 a new dawn for local learning methods the relentless improvement in speed of computers continues while some technical barriers to this progress have begun to emerge exploitation of parallelism has actually increased the rate of acceleration for many purposes especially in applied mathematical fields such as data mining interestingly new powerful hardware has been put to the task of running ever more baroque algorithms feedforward neural networks once trained over several days now train in minutes on affordable desktop hardware over time ever fancier algorithms have been fed to these machines boosting support vector machines random forests and most recently deep learning illustrate this trend another class of learning algorithms may also benefit from developments in hardware local learning methods typical of local methods are radial basis function rbf neural networks and k nearest neighbors k nn rbf neural networks were briefly popular in the heyday of neural networks the 1990s since they train much faster than the more popular feedforward neural networks k nn is often discussed in chapter 1 of machine learning books it is conceptually simple easy to implement and demonstrates the advantages and disadvantages of local techniques well local learning techniques usually have a large number of components each of which handles only a small fraction of the set of possible input cases the nice thing about this approach is that these local components largely do not need to coordinate with each other the complexity of the model comes from having a large number of such components to handle many different situations local learning techniques thus make training easy in the case of k nn one simply stores the training data for future reference little if any fitting is done during learning this gift comes with a price though local learning systems train very quickly but model execution is often rather slow this is because local models will either fire all of those local components or spend time figuring out which among them applies to any given situation local learning methods have largely fallen out of favor since 1 they are slow to predict outcomes for new cases and secondarily 2 their implementation requires retention of some or all of the training data and 2 this author wonders whether contemporary computer hardware may not present an opportunity for a resurgence among local methods local methods often perform well statistically and would help diversify model ensembles for users of more popular learning algoprithms analysts looking for that last measure of improvement might be well served by investigating this class of solutions local algorithms are among the easiest to code from scratch interested readers are directed to lazy learning edited by d aha isbn 13 978 0792345848 and nearest neighbor norms nn pattern classification techniques edited by b dasarathy isbn 13 978 0818689307 posted by will dwinnell at 08 27 no comments labels deep learning k nearest neighbor k nearest neighbors k nn lazy learning local learning nearest neighbor nearest neighbors neural network radial basis function rbf rbf neural network wednesday march 29 2017 data analytics summit iii at harrisburg university of science and technology harrisburg university of science and technology harrisburg pennsylvania has just finished hosting data analytics summit iii this is a multi day event featuring a mix of presenters from the private sector the government government related businesses and academia which spans research practice and more visionary big picture topics the theme was analytics applied case studies measuring impact and communicating results regrettably i was unable to attend this time because i was traveling for business but i was at data analytics summit ii which was held in december of 2015 if you haven t been harrisburg university of science and technology does a nice job hosting this event additionally so far the data analytics summit has been free of charge so there is the prospect of free food if you are a starving grad student the university has generously provided links to video of the presentations from the most recent summit http analyticssummit harrisburgu edu video links for the previous summit whose theme was unstructured data can be found at the bottom of my article unstructured data mining a primer apr 11 2016 over on icrunchdata https icrunchdata com unstructured data mining primer i encourage readers to explore this free resource posted by will dwinnell at 15 20 no comments labels conference data analytics summit ii data analytics summit iii free food harrisburg harrisburg university of science and technology pennsylvania summit video presentations friday march 17 2017 geographic distances a quick trip around the great circle recently i wanted to calculate the distance between locations on the earth finding a handy solution i thought readers might be interested in my situation location data included zip codes american postal codes also available to me is a look up table of the latitude and longitude of the geometric centroid of each zip code since the areas identified by zip codes are usually geographical small and making the close enough assumption that this planet is perfectly spherical trigonometry will allow distance calculations which are for most purposes precise enough given the latitude and longitude of cities a and b the following line of matlab code will calculate the distance between the two coordinates as the crow flies technically the great circle distance in kilometers distancekilometers round 111 12 acosd cosd longa longb cosd lata cosd latb sind lata sind latb note that latitude and longitude are expected as decimal degrees if your data is in degrees minutes seconds a quick conversion will be needed i ve checked this formula against a second source and quickly verified it using a few pairs of cities a new york b atlanta random on line reference 1202km lata 40 664274 longa 73 9385 latb 33 762909 longb 84 422675 distancekilometers round 111 12 acosd cosd longa longb cosd lata cosd latb sind lata sind latb distancekilometers 1202 a new york b los angeles random on line reference 3940km less than 0 5 difference 0 lata 40 664274 longa 73 9385 latb 34 019394 longb 118 410825 distancekilometers round 111 12 acosd cosd longa longb cosd lata cosd latb sind lata sind latb distancekilometers 3955 references how far is berlin by alan zeichick published in the sep 1991 issue of computer language magazine note that zeichick credits as his source an hp 27 scientific calculator from which he reverse engineered the formula above trigonometry demystified 2nd edition by stan gibilisco isbn 978 0 07 178024 7 posted by will dwinnell at 23 42 no comments labels distance between cities geographic distance geography global distance great circle great circle distance latitude location data longitude spatial statistics spherical trigonometry trigonometry friday march 18 2016 four books worth owning below are listed four books on statistics which i feel are worth owning they largely take a traditional statistics perspective as opposed to a machine learning data mining one with the exception of the statistical sleuth these are less like textbooks than guide books with information reflecting the experience and practical advice of their respective authors comparatively few of their pages are devoted to predictive modeling rather they cover a host of topics relevant to the statistical analyst sample size determination hypothesis testing assumptions sampling technique etc common errors in statistics by good and hardin statistical rules of thumb by van belle your statistical consultant by newton and rudestam the statistical sleuth by ramsey and schafer i have not given isbns since they vary by edition older editions of any of these will be valuable so consider saving money by skipping the latest edition a final clarification i am not giving a blanket endorsement to any of these books i completely disagree with a few of their ideas i see the value of such books in their use as paper advisers with different backgrounds and perspectives than my own posted by will dwinnell at 22 02 no comments labels books paper advisers statistics older posts home subscribe to posts atom about me will dwinnell i am a data scientist with more years of experience than i care to remember i ve worked in a variety of fields and used a wide array of tools but matlab is my tool of choice find me at http www linkedin com in predictor view my complete profile search data mining in matlab links rss feed for this web log data mining and predictive analytics the mathworks matlab vendor blog archive 2022 1 january 1 implied probabilities 2017 3 april 1 march 2 2016 3 march 3 2014 1 april 1 2013 2 november 1 january 1 2012 1 july 1 2010 5 december 1 september 1 february 3 2009 11 july 1 april 2 march 6 february 2 2008 11 november 1 april 2 march 8 2007 28 december 1 october 1 september 1 july 3 june 3 may 1 april 3 ...
|