Meta tags:
Headings (most frequently used words):
the, interpretation, about, simple, linear, regression, numerical, properties, intercept, variance, of, response, assumption, contents, formulation, and, computation, statistical, example, alternatives, see, also, references, external, links, expanded, formulas, relationship, with, sample, covariance, matrix, slope, correlation, unbiasedness, mean, predicted, confidence, intervals, line, fitting, without, term, single, regressor, normality, asymptotic,
Text of the page (most frequently used words):
the (307), displaystyle (119), and (97), #regression (83), widehat (78), beta (67), sum (63), bar (62), alpha (51), #linear (50), frac (50), left (48), right (48), model (35), for (35), this (34), that (33), data (29), var (29), hat (28), simple (27), line (27), edit (27), with (26), least (21), are (21), analysis (20), variable (20), aligned (20), correlation (19), varepsilon (19), statistics (18), sample (18), operatorname (18), squares (17), variance (17), confidence (17), distribution (16), from (15), estimator (15), mean (15), slope (15), intercept (15), function (14), statistical (14), points (14), can (14), values (14), sigma (14), random (13), error (13), begin (13), end (13), interpretation (13), test (12), overline (12), not (12), standard (11), one (11), term (11), assumption (11), value (11), response (11), limits (11), average (10), equations (10), residuals (10), point (10), example (10), given (10), about (9), may (9), time (9), coefficient (9), interval (9), when (9), where (9), cov (9), variables (9), ns_ (9), wikipedia (8), general (8), ordinary (8), errors (8), fit (8), estimators (8), plot (8), article (8), also (8), above (8), deviations (8), which (8), then (8), relationship (8), estimates (8), predicted (8), between (8), under (7), page (7), population (7), estimation (7), squared (7), unbiased (7), transformation (7), chart (7), polynomial (7), independent (7), used (7), see (7), measurement (7), more (7), here (7), set (7), dependent (7), will (7), sqrt (7), intervals (7), these (7), properties (7), toggle (6), fitting (6), series (6), likelihood (6), multivariate (6), covariance (6), generalized (6), bayesian (6), distance (6), matrix (6), formulas (6), through (6), because (6), single (6), equal (6), bmatrix (6), text (5), non (5), use (5), methods (5), mathematics (5), design (5), log (5), rank (5), equation (5), effects (5), median (5), normality (5), theorem (5), parameter (5), experiment (5), absolute (5), center (5), all (5), other (5), 5pt (5), quantile (5), numerical (5), large (5), such (5), since (5), each (5), rho (5), 1ex (5), delta (5), table (4), contents (4), search (4), additional (4), terms (4), was (4), retrieved (4), exponential (4), information (4), system (4), models (4), product (4), limit (4), box (4), partial (4), statistic (4), normal (4), degrees (4), freedom (4), anova (4), mixed (4), inference (4), student (4), power (4), scale (4), normalization (4), how (4), links (4), tools (4), 1548 (4), called (4), case (4), mass (4), 6pt (4), without (4), weighted (4), vertical (4), has (4), should (4), but (4), underlying (4), 8pt (4), would (4), coefficients (4), number (4), shown (4), level (4), changes (4), distributed (4), have (4), zero (4), expression (4), some (4), possible (4), formulation (4), known (4), angle (4), expanded (4), hide (4), move (4), sidebar (4), subsection (4), view (3), using (3), commons (3), articles (3), description (3), parametric (3), index (3), moving (3), approach (3), trend (3), portal (3), control (3), first (3), survival (3), spectral (3), vector (3), cross (3), tests (3), factor (3), principal (3), logistic (3), robust (3), nonparametric (3), nonlinear (3), adaptive (3), validation (3), determination (3), pearson (3), moment (3), ordered (3), ratio (3), prediction (3), method (3), moments (3), estimating (3), optimal (3), family (3), empirical (3), order (3), sampling (3), study (3), size (3), scatter (3), central (3), geometric (3), university (3), calculate (3), maths (3), pdf (3), involving (3), form (3), gives (3), appropriate (3), origin (3), ols (3), allows (3), best (3), related (3), section (3), does (3), into (3), its (3), minimizing (3), pairs (3), there (3), explanatory (3), parameters (3), alternatives (3), their (3), qquad (3), 931 (3), 0532 (3), 58498 (3), 5439 (3), 2453 (3), following (3), derived (3), law (3), fixed (3), account (3), normally (3), same (3), make (3), means (3), theta (3), get (3), unobserved (3), computation (3), fitted (3), probit (3), main (3), topic (2), languages (2), contact (2), privacy (2), policy (2), wikimedia (2), last (2), categories (2), clarification (2), 2015 (2), short (2), wikidata (2), php (2), title (2), forecasts (2), decomposition (2), smoothing (2), process (2), engineering (2), studies (2), clinical (2), proportional (2), density (2), frequency (2), domain (2), autoregressive (2), specific (2), structural (2), seasonal (2), adjustment (2), stationarity (2), distributions (2), cluster (2), components (2), contingency (2), categorical (2), poisson (2), regressions (2), binomial (2), isotonic (2), semiparametric (2), maximum (2), posterior (2), prior (2), probability (2), van (2), alternative (2), way (2), lehmann (2), goodness (2), score (2), most (2), bootstrap (2), minimum (2), location (2), shape (2), natural (2), designs (2), trial (2), randomized (2), controlled (2), survey (2), missing (2), reduction (2), cleaning (2), scaling (2), transform (2), run (2), dispersion (2), deviation (2), range (2), arithmetic (2), external (2), isbn (2), 3rd (2), applied (2), samples (2), new (2), 2024 (2), jan (2), muthukrishnan (2), behind (2), pmid (2), oclc (2), issn (2), doi (2), 227 (2), introduction (2), references (2), proofs (2), segmented (2), uncorrected (2), substituting (2), place (2), assumed (2), regressor (2), invariant (2), equivalent (2), units (2), axis (2), deming (2), changing (2), outliers (2), constructing (2), straight (2), written (2), elsewhere (2), titles (2), while (2), found (2), talk (2), total (2), finds (2), two (2), dimensional (2), coordinates (2), than (2), whose (2), slr (2), only (2), measured (2), 9946 (2), might (2), calculated (2), graph (2), 272 (2), 062 (2), 5762 (2), 1539 (2), 63185 (2), height (2), women (2), quadratic (2), second (2), approximately (2), previous (2), replaced (2), numbers (2), asymptotic (2), bands (2), 859 (2), 817 (2), give (2), okun (2), unemployment (2), gdp (2), growth (2), construct (2), independently (2), justified (2), estimated (2), follows (2), fact (2), further (2), defined (2), estimate (2), consider (2), drawn (2), words (2), true (2), unbiasedness (2), includes (2), what (2), imagine (2), discrete (2), arctan (2), therefore (2), tan (2), tangent (2), small (2), notation (2), over (2), passes (2), item (2), original (2), leq (2), solution (2), respectively (2), expanding (2), efficient (2), version (2), argmin (2), goal (2), call (2), residual (2), logit (2), multinomial (2), appearance (2), upload (2), file (2), history (2), read (2), create (2), donate (2), menu (2), add, mobile, cookie, statement, developers, code, conduct, legal, safety, contacts, disclaimers, available, apply, site, you, agree, registered, trademark, profit, organization, foundation, inc, creative, attribution, sharealike, license, rendered, parsoid, edited, february, 2026, utc, hidden, excerpts, needing, october, matches, curve, https, org, simple_linear_regression, oldid, 1336583492, associative, causal, econometric, historical, naïve, quantitative, forecasting, wikiproject, category, kriging, geostatistics, geographic, environmental, cartography, spatial, psychometrics, official, national, accounts, jurimetrics, econometrics, demography, crime, census, actuarial, science, social, identification, reliability, quality, probabilistic, chemometrics, medical, epidemiology, trials, bioinformatics, biostatistics, applications, nelson, aalen, hazard, hitting, accelerated, failure, aft, hazards, kaplan, meier, whittle, wavelet, fourier, autoregression, conditional, heteroskedasticity, arch, arima, jenkins, arma, xcf, pacf, autocorrelation, acf, breusch, godfrey, durbin, watson, ljung, johansen, dickey, fuller, granger, causality, break, cointegration, elliptical, classification, discriminant, canonical, manova, cochran, mantel, haenszel, mcnemar, graphical, cohen, kappa, partition, bernoulli, families, homoscedasticity, heteroscedasticity, predictors, template, splines, mars, simultaneous, confounding, bayes, credible, der, waerden, jonckheere, terpstra, friedman, kruskal, wallis, mann, whitney, hodges, signed, wilcoxon, sign, bic, aic, selection, shapiro, wilk, jarque, bera, lilliefors, anderson, darling, kolmogorov, smirnov, chi, wald, lagrange, multiplier, multiple, comparisons, randomization, permutation, uniformly, powerful, tails, testing, hypotheses, jackknife, resampling, tolerance, pivot, plug, scheffé, rao, blackwellization, frequentist, robustness, asymptotics, divergence, efficiency, loss, decision, functional, sufficiency, completeness, monotone, space, specification, theory, quasi, sectional, cohort, observational, down, stochastic, approximation, scientific, assignment, interaction, factorial, blocking, experiments, questionnaire, opinion, poll, stratified, methodology, replication, effect, collection, detrending, differencing, preprocessing, component, dimensionality, truncation, winsorizing, outlier, unit, min, max, standardization, feature, fisher, anscombe, stabilizing, yeo, johnson, cox, transformations, processing, ecdf, heatmap, violin, stem, leaf, display, radar, pie, histogram, forest, fan, correlogram, biplot, graphics, spearman, kendall, dependence, grouped, summary, tables, count, skewness, kurtosis, percentile, interquartile, variation, mode, lehmer, heinz, heronian, harmonic, cubic, contraharmonic, continuous, descriptive, outline, robert, nau, duke, wolfram, mathworld, explanation, casella, berger, 2002, 2nd, edition, cengage, 558, 559, 978, 534, 24312, draper, smith, 1998, john, wiley, 471, 17082, valliant, richard, jill, dever, frauke, kreuter, practical, designing, weighting, york, springer, 2013, numeracy, academic, skills, kit, newcastle, class, gowri, jun, 2018, kenney, keeping, 1962, princeton, nostrand, 252, 285, altman, naomi, krzywinski, martin, 1000, 261269711, s2cid, 26824102, 5912005539, 7091, 1038, nmeth, 3627, 999, nature, zou, tuncali, silverman, 2003, 12773666, 110941167, 0033, 8419, 1148, radiol, 2273011499, 617, radiology, lane, david, 462, columbia, 2016, seltman, howard, 2008, experimental, newey, west, derivation, multidimensional, refer, bias, demonstrates, away, affects, sometimes, force, pass, simplifies, both, altered, major, leads, different, orthogonal, perpendicular, resistance, several, exist, considering, editors, believe, holds, described, unrelated, moved, 2019, relevant, discussion, disambiguation, drafted, broad, concept, primary, excerpt, fits, unlike, really, instance, separate, could, potentially, return, lead, attempts, include, chooses, slopes, determined, theil, sen, contains, biased, due, dilution, calculating, 975, thus, 1604, lines, quantities, hand, calculations, started, finding, five, sums, 5544, 2916, 136, 2618, 3489, 5211, 3961, 129, 9420, 2400, 4888, 8064, 124, 4576, 1684, 4637, 6100, 119, 1750, 0625, 4393, 0384, 114, 6644, 9929, 4156, 3809, 109, 5990, 8900, 3982, 8721, 106, 0248, 8224, 3756, 4641, 101, 1285, 7225, 3591, 6049, 6859, 6569, 3430, 4449, 7120, 5600, 3271, 8400, 8040, 4649, 3118, 1056, 5520, 4025, 2968, 0704, 8096, 3104, 2821, 7344, 6800, 2500, 2725, 8841, 7487, 1609, masses, american, age, although, argues, instead, states, dataset, enough, become, applicable, remain, valid, exception, occasionally, fraction, change, alter, results, appreciably, turns, cdot, represent, graphically, around, proceed, carefully, joint, band, hyperbolic, idea, likely, similarly, scriptstyle, gamma, sim, itself, proportionally, textstyle, latter, observations, sufficiently, classic, relies, either, allow, however, those, tell, precise, much, vary, specified, were, devised, plausible, repeated, very, times, 4pt, unknown, additionally, earlier, demonstrate, simplification, identity, simplified, 2x_, context, every, observation, say, mid, equiv, formalize, assertion, must, define, framework, corresponding, generated, plus, themselves, definition, requires, based, assuming, validity, evaluate, assumptions, discussed, needed, inhomogeneity, variances, extreme, self, evident, 401, uncorrelated, whether, meaning, goes, forced, framing, actually, type, issue, defines, our, had, expectation, positive, think, just, uniform, notice, constant, upfront, depend, deriving, showing, makes, connects, important, position, affect, connecting, multiplying, members, summation, numerator, thereby, details, concise, formula, generalizing, write, horizontal, indicate, shows, followup, expect, closer, phenomenon, toward, standardized, expressions, yields, reformulated, elements, 2ex, solved, directly, stand, alone, resultant, algebraically, ones, paragraph, below, proof, calculation, defining, respect, respective, introduced, derive, arguments, denoted, objective, solve, minimization, problem, find, provide, sense, mentioned, understood, minimizes, differences, actual, any, candidate, describes, hold, exactly, largely, suppose, observe, them, describe, common, stipulation, accuracy, corrected, concerns, conventionally, accurately, predicts, adjective, refers, outcome, predictor, cartesian, coordinate, gauss, markov, studentized, background, iteratively, reweighted, regularized, ridge, negative, local, multilevel, binary, choice, part, presumed, rate, macroeconomics, free, encyclopedia, projects, printable, download, print, export, switch, legacy, parser, shortened, url, cite, permanent, link, actions, english, українська, português, 한국어, 日本語, bahasa, indonesia, עברית, فارسی, eesti, ελληνικά, deutsch, català, беларуская, العربية, top, personal, special, pages, recent, community, learn, help, contribute, current, events, navigation, jump, content,
Text of the page (random words):
able x as we just defined it numerical properties edit the regression line goes through the center of mass point x y displaystyle bar x bar y if the model includes an intercept term i e not forced through the origin the sum of the residuals is zero if the model includes an intercept term i 1 n ε i 0 displaystyle sum _ i 1 n widehat varepsilon _ i 0 the residuals and x values are uncorrelated whether or not there is an intercept term in the model meaning i 1 n x i ε i 0 displaystyle sum _ i 1 n x_ i widehat varepsilon _ i 0 the relationship between ρ x y displaystyle rho _ xy the correlation coefficient for the population and the population variances of y displaystyle y σ y 2 displaystyle sigma _ y 2 and the error term of ε displaystyle varepsilon σ ε 2 displaystyle sigma _ varepsilon 2 is 10 401 σ ε 2 1 ρ x y 2 σ y 2 displaystyle sigma _ varepsilon 2 1 rho _ xy 2 sigma _ y 2 for extreme values of ρ x y displaystyle rho _ xy this is self evident since when ρ x y 0 displaystyle rho _ xy 0 then σ ε 2 σ y 2 displaystyle sigma _ varepsilon 2 sigma _ y 2 and when ρ x y 1 displaystyle rho _ xy 1 then σ ε 2 0 displaystyle sigma _ varepsilon 2 0 statistical properties edit description of the statistical properties of estimators from the simple linear regression estimates requires the use of a statistical model the following is based on assuming the validity of a model under which the estimates are optimal it is also possible to evaluate the properties under other assumptions such as inhomogeneity but this is discussed elsewhere clarification needed unbiasedness edit the estimators α displaystyle widehat alpha and β displaystyle widehat beta are unbiased to formalize this assertion we must define a framework in which these estimators are random variables we consider the residuals ε i as random variables drawn independently from some distribution with mean zero in other words for each value of x the corresponding value of y is generated as a mean response α βx plus an additional random variable ε called the error term equal to zero on average under such interpretation the least squares estimators α displaystyle widehat alpha and β displaystyle widehat beta will themselves be random variables whose means will equal the true values α and β this is the definition of an unbiased estimator variance of the mean response edit since the data in this context is defined to be x y pairs for every observation the mean response at a given value of x say x d is an estimate of the mean of the y values in the population at the x value of x d that is e y x d y d displaystyle hat e y mid x_ d equiv hat y _ d the variance of the mean response is given by 11 var α β x d var α var β x d 2 2 x d cov α β displaystyle operatorname var left hat alpha hat beta x_ d right operatorname var left hat alpha right left operatorname var hat beta right x_ d 2 2x_ d operatorname cov left hat alpha hat beta right this expression can be simplified to var α β x d σ 2 1 m x d x 2 x i x 2 displaystyle operatorname var left hat alpha hat beta x_ d right sigma 2 left frac 1 m frac left x_ d bar x right 2 sum x_ i bar x 2 right where m is the number of data points to demonstrate this simplification one can make use of the identity i x i x 2 i x i 2 1 m i x i 2 displaystyle sum _ i x_ i bar x 2 sum _ i x_ i 2 frac 1 m left sum _ i x_ i right 2 variance of the predicted response edit further information prediction interval the predicted response distribution is the predicted distribution of the residuals at the given point x d so the variance is given by var y d α β x d var y d var α β x d 2 cov y d α β x d var y d var α β x d displaystyle begin aligned operatorname var left y_ d left hat alpha hat beta x_ d right right operatorname var y_ d operatorname var left hat alpha hat beta x_ d right 2 operatorname cov left y_ d left hat alpha hat beta x_ d right right operatorname var y_ d operatorname var left hat alpha hat beta x_ d right end aligned the second line follows from the fact that cov y d α β x d displaystyle operatorname cov left y_ d left hat alpha hat beta x_ d right right is zero because the new prediction point is independent of the data used to fit the model additionally the term var α β x d displaystyle operatorname var left hat alpha hat beta x_ d right was calculated earlier for the mean response since var y d σ 2 displaystyle operatorname var y_ d sigma 2 a fixed but unknown parameter that can be estimated the variance of the predicted response is given by var y d α β x d σ 2 σ 2 1 m x d x 2 x i x 2 σ 2 1 1 m x d x 2 x i x 2 displaystyle begin aligned operatorname var left y_ d left hat alpha hat beta x_ d right right sigma 2 sigma 2 left frac 1 m frac left x_ d bar x right 2 sum x_ i bar x 2 right 4pt sigma 2 left 1 frac 1 m frac x_ d bar x 2 sum x_ i bar x 2 right end aligned confidence intervals edit the formulas given in the previous section allow one to calculate the point estimates of α and β that is the coefficients of the regression line for the given set of data however those formulas do not tell us how precise the estimates are i e how much the estimators α displaystyle widehat alpha and β displaystyle widehat beta vary from sample to sample for the specified sample size confidence intervals were devised to give a plausible set of values to the estimates one might have if one repeated the experiment a very large number of times the standard method of constructing confidence intervals for linear regression coefficients relies on the normality assumption which is justified if either the errors in the regression are normally distributed the so called classic regression assumption or the number of observations n is sufficiently large in which case the estimator is approximately normally distributed the latter case is justified by the central limit theorem normality assumption edit under the first assumption above that of the normality of the error terms the estimator of the slope coefficient will itself be normally distributed with mean β and variance σ 2 i x i x 2 textstyle sigma 2 left sum _ i x_ i bar x 2 right where σ 2 is the variance of the error terms see proofs involving ordinary least squares at the same time the sum of squared residuals q is distributed proportionally to χ 2 with n 2 degrees of freedom and independently from β displaystyle widehat beta this allows us to construct a t value t β β s β t n 2 displaystyle t frac widehat beta beta s_ widehat beta sim t_ n 2 where s β 1 n 2 i 1 n ε i 2 i 1 n x i x 2 displaystyle s_ widehat beta sqrt frac frac 1 n 2 sum _ i 1 n widehat varepsilon _ i 2 sum _ i 1 n x_ i bar x 2 is the unbiased standard error estimator of the estimator β displaystyle widehat beta this t value has a student s t distribution with n 2 degrees of freedom using it we can construct a confidence interval for β β β s β t n 2 β s β t n 2 displaystyle beta in left widehat beta s_ widehat beta t_ n 2 widehat beta s_ widehat beta t_ n 2 right at confidence level 1 γ where t n 2 displaystyle t_ n 2 is the 1 γ 2 th displaystyle scriptstyle left 1 frac gamma 2 right text th quantile of the t n 2 distribution for example if γ 0 05 then the confidence level is 95 similarly the confidence interval for the intercept coefficient α is given by α α s α t n 2 α s α t n 2 displaystyle alpha in left widehat alpha s_ widehat alpha t_ n 2 widehat alpha s_ widehat alpha t_ n 2 right at confidence level 1 γ where s α s β 1 n i 1 n x i 2 1 n n 2 i 1 n ε i 2 i 1 n x i 2 i 1 n x i x 2 displaystyle s_ widehat alpha s_ widehat beta sqrt frac 1 n sum _ i 1 n x_ i 2 sqrt frac 1 n n 2 left sum _ i 1 n widehat varepsilon _ i 2 right frac sum _ i 1 n x_ i 2 sum _ i 1 n x_ i bar x 2 the us changes in unemployment gdp growth regression with the 95 confidence bands the confidence intervals for α and β give us the general idea where these regression coefficients are most likely to be for example in the okun s law regression shown here the point estimates are α 0 859 β 1 817 displaystyle widehat alpha 0 859 qquad widehat beta 1 817 the 95 confidence intervals for these estimates are α 0 76 0 96 β 2 06 1 58 displaystyle alpha in left 0 76 0 96 right qquad beta in left 2 06 1 58 right in order to represent this information graphically in the form of the confidence bands around the regression line one has to proceed carefully and account for the joint distribution of the estimators it can be shown 12 that at confidence level 1 γ the confidence band has hyperbolic form given by the equation α β ξ α β ξ t n 2 1 n 2 ε i 2 1 n ξ x 2 x i x 2 displaystyle alpha beta xi in left widehat alpha widehat beta xi pm t_ n 2 sqrt left frac 1 n 2 sum widehat varepsilon _ i 2 right cdot left frac 1 n frac xi bar x 2 sum x_ i bar x 2 right right when the model assumed the intercept is fixed and equal to 0 α 0 displaystyle alpha 0 the standard error of the slope turns into s β 1 n 1 i 1 n ε i 2 i 1 n x i 2 displaystyle s_ widehat beta sqrt frac 1 n 1 frac sum _ i 1 n widehat varepsilon _ i 2 sum _ i 1 n x_ i 2 with ε i y i y i displaystyle hat varepsilon _ i y_ i hat y _ i asymptotic assumption edit the alternative second assumption states that when the number of points in the dataset is large enough the law of large numbers and the central limit theorem become applicable and then the distribution of the estimators is approximately normal under this assumption all formulas derived in the previous section remain valid with the only exception that the quantile t n 2 of student s t distribution is replaced with the quantile q of the standard normal distribution occasionally the fraction 1 n 2 is replaced with 1 n when n is large such a change does not alter the results appreciably numerical example edit see also ordinary least squares example and linear least squares example this data set gives average masses for women as a function of their height in a sample of american women of age 30 39 although the ols article argues that it would be more appropriate to run a quadratic regression for this data the simple linear regression model is applied here instead height m x i 1 47 1 50 1 52 1 55 1 57 1 60 1 63 1 65 1 68 1 70 1 73 1 75 1 78 1 80 1 83 mass kg y i 52 21 53 12 54 48 55 84 57 20 58 57 59 93 61 29 63 11 64 47 66 28 68 10 69 92 72 19 74 46 i displaystyle i x i displaystyle x_ i y i displaystyle y_ i x i 2 displaystyle x_ i 2 x i y i displaystyle x_ i y_ i y i 2 displaystyle y_ i 2 1 1 47 52 21 2 1609 76 7487 2725 8841 2 1 50 53 12 2 2500 79 6800 2821 7344 3 1 52 54 48 2 3104 82 8096 2968 0704 4 1 55 55 84 2 4025 86 5520 3118 1056 5 1 57 57 20 2 4649 89 8040 3271 8400 6 1 60 58 57 2 5600 93 7120 3430 4449 7 1 63 59 93 2 6569 97 6859 3591 6049 8 1 65 61 29 2 7225 101 1285 3756 4641 9 1 68 63 11 2 8224 106 0248 3982 8721 10 1 70 64 47 2 8900 109 5990 4156 3809 11 1 73 66 28 2 9929 114 6644 4393 0384 12 1 75 68 10 3 0625 119 1750 4637 6100 13 1 78 69 92 3 1684 124 4576 4888 8064 14 1 80 72 19 3 2400 129 9420 5211 3961 15 1 83 74 46 3 3489 136 2618 5544 2916 σ displaystyle sigma 24 76 931 17 41 0532 1548 2453 58498 5439 there are n 15 points in this data set hand calculations would be started by finding the following five sums s x i x i 24 76 s y i y i 931 17 s x x i x i 2 41 0532 s y y i y i 2 58498 5439 s x y i x i y i 1548 2453 displaystyle begin aligned s_ x sum _ i x_ i 24 76 qquad s_ y sum _ i y_ i 931 17 5pt s_ xx sum _ i x_ i 2 41 0532 s_ yy sum _ i y_ i 2 58498 5439 5pt s_ xy sum _ i x_ i y_ i 1548 2453 end aligned these quantities would be used to calculate the estimates of the regression coefficients and their standard errors β n s x y s x s y n s x x s x 2 61 272 α 1 n s y β 1 n s x 39 062 s ε 2 1 n n 2 n s y y s y 2 β 2 n s x x s x 2 0 5762 s β 2 n s ε 2 n s x x s x 2 3 1539 s α 2 s β 2 1 n s x x 8 63185 displaystyle begin aligned widehat beta frac ns_ xy s_ x s_ y ns_ xx s_ x 2 61 272 8pt widehat alpha frac 1 n s_ y widehat beta frac 1 n s_ x 39 062 8pt s_ varepsilon 2 frac 1 n n 2 left ns_ yy s_ y 2 widehat beta 2 ns_ xx s_ x 2 right 0 5762 8pt s_ widehat beta 2 frac ns_ varepsilon 2 ns_ xx s_ x 2 3 1539 8pt s_ widehat alpha 2 s_ widehat beta 2 frac 1 n s_ xx 8 63185 end aligned graph of points and linear least squares lines in the simple linear regression numerical example the 0 975 quantile of student s t distribution with 13 degrees of freedom is t 13 2 1604 and thus the 95 confidence intervals for α and β are α α t 13 s α 45 4 32 7 β β t 13 s β 57 4 65 1 displaystyle begin aligned alpha in widehat alpha mp t_ 13 s_ widehat alpha 45 4 32 7 5pt beta in widehat beta mp t_ 13 s_ widehat beta 57 4 65 1 end aligned the product moment correlation coefficient might also be calculated r n s x y s x s y n s x x s x 2 n s y y s y 2 0 9946 displaystyle widehat r frac ns_ xy s_ x s_ y sqrt ns_ xx s_ x 2 ns_ yy s_ y 2 0 9946 alternatives edit calculating the parameters of a linear model by minimizing the squared error in slr there is an underlying assumption that only the dependent variable contains measurement error if the explanatory variable is also measured with error then simple regression is not appropriate for estimating the underlying relationship because it will be biased due to regression dilution other estimation methods that can be used in place of ordinary least squares include least absolute deviations minimizing the sum of absolute values of residuals and the theil sen estimator which chooses a line whose slope is the median of the slopes determined by pairs of sample points deming regression total least squares also finds a line that fits a set of two dimensional sample points but unlike ordinary least squares least absolute deviations and median slope regression it is not really an instance of simple linear regression because it does not separate the coordinates into one dependent and one independent variable and could potentially return a vertical line as its fit can lead to a model that attempts to fit the outliers more than the data line fitting edit this section is an excerpt from line fitting edit this page is a primary topic and an article should be written about it one or more editors believe it holds the title of a broad concept article the article may be written here or drafted elsewhere first related titles should be described here while unrelated titles should be moved to simple linear regression disambiguation relevant discussion may be found on the talk page may 2019 line fitting is the process of constructing a straight line that has the best fit to a series of data points several methods exist considering vertical distance simple linear regression resistance to outliers robust simple linear regression perpendicular distance orthogonal regression this is not scale invariant i e changing the measurement units leads to a different line weighted geometric distance deming regression scale invariant approach major axis regression this allows for measurement error in both variables and gives an equivalent equation if the measurement units are altere...
|