Meta tags:
Headings (most frequently used words):
the, interpretation, about, simple, linear, regression, numerical, properties, intercept, variance, of, response, assumption, contents, formulation, and, computation, statistical, example, alternatives, see, also, references, external, links, expanded, formulas, relationship, with, sample, covariance, matrix, slope, correlation, unbiasedness, mean, predicted, confidence, intervals, line, fitting, without, term, single, regressor, normality, asymptotic,
Text of the page (most frequently used words):
the (307), displaystyle (119), and (97), #regression (83), widehat (78), beta (67), sum (63), bar (62), alpha (51), #linear (50), frac (50), left (48), right (48), model (35), for (35), this (34), that (33), data (29), var (29), hat (28), simple (27), line (27), edit (27), with (26), least (21), are (21), analysis (20), variable (20), aligned (20), correlation (19), varepsilon (19), statistics (18), sample (18), operatorname (18), squares (17), variance (17), confidence (17), distribution (16), from (15), estimator (15), mean (15), slope (15), intercept (15), function (14), statistical (14), points (14), can (14), values (14), sigma (14), random (13), error (13), begin (13), end (13), interpretation (13), test (12), overline (12), not (12), standard (11), one (11), term (11), assumption (11), value (11), response (11), limits (11), average (10), equations (10), residuals (10), point (10), example (10), given (10), about (9), may (9), time (9), coefficient (9), interval (9), when (9), where (9), cov (9), variables (9), ns_ (9), wikipedia (8), general (8), ordinary (8), errors (8), fit (8), estimators (8), plot (8), article (8), also (8), above (8), deviations (8), which (8), then (8), relationship (8), estimates (8), predicted (8), between (8), under (7), page (7), population (7), estimation (7), squared (7), unbiased (7), transformation (7), chart (7), polynomial (7), independent (7), used (7), see (7), measurement (7), more (7), here (7), set (7), dependent (7), will (7), sqrt (7), intervals (7), these (7), properties (7), toggle (6), fitting (6), series (6), likelihood (6), multivariate (6), covariance (6), generalized (6), bayesian (6), distance (6), matrix (6), formulas (6), through (6), because (6), single (6), equal (6), bmatrix (6), text (5), non (5), use (5), methods (5), mathematics (5), design (5), log (5), rank (5), equation (5), effects (5), median (5), normality (5), theorem (5), parameter (5), experiment (5), absolute (5), center (5), all (5), other (5), 5pt (5), quantile (5), numerical (5), large (5), such (5), since (5), each (5), rho (5), 1ex (5), delta (5), table (4), contents (4), search (4), additional (4), terms (4), was (4), retrieved (4), exponential (4), information (4), system (4), models (4), product (4), limit (4), box (4), partial (4), statistic (4), normal (4), degrees (4), freedom (4), anova (4), mixed (4), inference (4), student (4), power (4), scale (4), normalization (4), how (4), links (4), tools (4), 1548 (4), called (4), case (4), mass (4), 6pt (4), without (4), weighted (4), vertical (4), has (4), should (4), but (4), underlying (4), 8pt (4), would (4), coefficients (4), number (4), shown (4), level (4), changes (4), distributed (4), have (4), zero (4), expression (4), some (4), possible (4), formulation (4), known (4), angle (4), expanded (4), hide (4), move (4), sidebar (4), subsection (4), view (3), using (3), commons (3), articles (3), description (3), parametric (3), index (3), moving (3), approach (3), trend (3), portal (3), control (3), first (3), survival (3), spectral (3), vector (3), cross (3), tests (3), factor (3), principal (3), logistic (3), robust (3), nonparametric (3), nonlinear (3), adaptive (3), validation (3), determination (3), pearson (3), moment (3), ordered (3), ratio (3), prediction (3), method (3), moments (3), estimating (3), optimal (3), family (3), empirical (3), order (3), sampling (3), study (3), size (3), scatter (3), central (3), geometric (3), university (3), calculate (3), maths (3), pdf (3), involving (3), form (3), gives (3), appropriate (3), origin (3), ols (3), allows (3), best (3), related (3), section (3), does (3), into (3), its (3), minimizing (3), pairs (3), there (3), explanatory (3), parameters (3), alternatives (3), their (3), qquad (3), 931 (3), 0532 (3), 58498 (3), 5439 (3), 2453 (3), following (3), derived (3), law (3), fixed (3), account (3), normally (3), same (3), make (3), means (3), theta (3), get (3), unobserved (3), computation (3), fitted (3), probit (3), main (3), topic (2), languages (2), contact (2), privacy (2), policy (2), wikimedia (2), last (2), categories (2), clarification (2), 2015 (2), short (2), wikidata (2), php (2), title (2), forecasts (2), decomposition (2), smoothing (2), process (2), engineering (2), studies (2), clinical (2), proportional (2), density (2), frequency (2), domain (2), autoregressive (2), specific (2), structural (2), seasonal (2), adjustment (2), stationarity (2), distributions (2), cluster (2), components (2), contingency (2), categorical (2), poisson (2), regressions (2), binomial (2), isotonic (2), semiparametric (2), maximum (2), posterior (2), prior (2), probability (2), van (2), alternative (2), way (2), lehmann (2), goodness (2), score (2), most (2), bootstrap (2), minimum (2), location (2), shape (2), natural (2), designs (2), trial (2), randomized (2), controlled (2), survey (2), missing (2), reduction (2), cleaning (2), scaling (2), transform (2), run (2), dispersion (2), deviation (2), range (2), arithmetic (2), external (2), isbn (2), 3rd (2), applied (2), samples (2), new (2), 2024 (2), jan (2), muthukrishnan (2), behind (2), pmid (2), oclc (2), issn (2), doi (2), 227 (2), introduction (2), references (2), proofs (2), segmented (2), uncorrected (2), substituting (2), place (2), assumed (2), regressor (2), invariant (2), equivalent (2), units (2), axis (2), deming (2), changing (2), outliers (2), constructing (2), straight (2), written (2), elsewhere (2), titles (2), while (2), found (2), talk (2), total (2), finds (2), two (2), dimensional (2), coordinates (2), than (2), whose (2), slr (2), only (2), measured (2), 9946 (2), might (2), calculated (2), graph (2), 272 (2), 062 (2), 5762 (2), 1539 (2), 63185 (2), height (2), women (2), quadratic (2), second (2), approximately (2), previous (2), replaced (2), numbers (2), asymptotic (2), bands (2), 859 (2), 817 (2), give (2), okun (2), unemployment (2), gdp (2), growth (2), construct (2), independently (2), justified (2), estimated (2), follows (2), fact (2), further (2), defined (2), estimate (2), consider (2), drawn (2), words (2), true (2), unbiasedness (2), includes (2), what (2), imagine (2), discrete (2), arctan (2), therefore (2), tan (2), tangent (2), small (2), notation (2), over (2), passes (2), item (2), original (2), leq (2), solution (2), respectively (2), expanding (2), efficient (2), version (2), argmin (2), goal (2), call (2), residual (2), logit (2), multinomial (2), appearance (2), upload (2), file (2), history (2), read (2), create (2), donate (2), menu (2), add, mobile, cookie, statement, developers, code, conduct, legal, safety, contacts, disclaimers, available, apply, site, you, agree, registered, trademark, profit, organization, foundation, inc, creative, attribution, sharealike, license, rendered, parsoid, edited, february, 2026, utc, hidden, excerpts, needing, october, matches, curve, https, org, simple_linear_regression, oldid, 1336583492, associative, causal, econometric, historical, naïve, quantitative, forecasting, wikiproject, category, kriging, geostatistics, geographic, environmental, cartography, spatial, psychometrics, official, national, accounts, jurimetrics, econometrics, demography, crime, census, actuarial, science, social, identification, reliability, quality, probabilistic, chemometrics, medical, epidemiology, trials, bioinformatics, biostatistics, applications, nelson, aalen, hazard, hitting, accelerated, failure, aft, hazards, kaplan, meier, whittle, wavelet, fourier, autoregression, conditional, heteroskedasticity, arch, arima, jenkins, arma, xcf, pacf, autocorrelation, acf, breusch, godfrey, durbin, watson, ljung, johansen, dickey, fuller, granger, causality, break, cointegration, elliptical, classification, discriminant, canonical, manova, cochran, mantel, haenszel, mcnemar, graphical, cohen, kappa, partition, bernoulli, families, homoscedasticity, heteroscedasticity, predictors, template, splines, mars, simultaneous, confounding, bayes, credible, der, waerden, jonckheere, terpstra, friedman, kruskal, wallis, mann, whitney, hodges, signed, wilcoxon, sign, bic, aic, selection, shapiro, wilk, jarque, bera, lilliefors, anderson, darling, kolmogorov, smirnov, chi, wald, lagrange, multiplier, multiple, comparisons, randomization, permutation, uniformly, powerful, tails, testing, hypotheses, jackknife, resampling, tolerance, pivot, plug, scheffé, rao, blackwellization, frequentist, robustness, asymptotics, divergence, efficiency, loss, decision, functional, sufficiency, completeness, monotone, space, specification, theory, quasi, sectional, cohort, observational, down, stochastic, approximation, scientific, assignment, interaction, factorial, blocking, experiments, questionnaire, opinion, poll, stratified, methodology, replication, effect, collection, detrending, differencing, preprocessing, component, dimensionality, truncation, winsorizing, outlier, unit, min, max, standardization, feature, fisher, anscombe, stabilizing, yeo, johnson, cox, transformations, processing, ecdf, heatmap, violin, stem, leaf, display, radar, pie, histogram, forest, fan, correlogram, biplot, graphics, spearman, kendall, dependence, grouped, summary, tables, count, skewness, kurtosis, percentile, interquartile, variation, mode, lehmer, heinz, heronian, harmonic, cubic, contraharmonic, continuous, descriptive, outline, robert, nau, duke, wolfram, mathworld, explanation, casella, berger, 2002, 2nd, edition, cengage, 558, 559, 978, 534, 24312, draper, smith, 1998, john, wiley, 471, 17082, valliant, richard, jill, dever, frauke, kreuter, practical, designing, weighting, york, springer, 2013, numeracy, academic, skills, kit, newcastle, class, gowri, jun, 2018, kenney, keeping, 1962, princeton, nostrand, 252, 285, altman, naomi, krzywinski, martin, 1000, 261269711, s2cid, 26824102, 5912005539, 7091, 1038, nmeth, 3627, 999, nature, zou, tuncali, silverman, 2003, 12773666, 110941167, 0033, 8419, 1148, radiol, 2273011499, 617, radiology, lane, david, 462, columbia, 2016, seltman, howard, 2008, experimental, newey, west, derivation, multidimensional, refer, bias, demonstrates, away, affects, sometimes, force, pass, simplifies, both, altered, major, leads, different, orthogonal, perpendicular, resistance, several, exist, considering, editors, believe, holds, described, unrelated, moved, 2019, relevant, discussion, disambiguation, drafted, broad, concept, primary, excerpt, fits, unlike, really, instance, separate, could, potentially, return, lead, attempts, include, chooses, slopes, determined, theil, sen, contains, biased, due, dilution, calculating, 975, thus, 1604, lines, quantities, hand, calculations, started, finding, five, sums, 5544, 2916, 136, 2618, 3489, 5211, 3961, 129, 9420, 2400, 4888, 8064, 124, 4576, 1684, 4637, 6100, 119, 1750, 0625, 4393, 0384, 114, 6644, 9929, 4156, 3809, 109, 5990, 8900, 3982, 8721, 106, 0248, 8224, 3756, 4641, 101, 1285, 7225, 3591, 6049, 6859, 6569, 3430, 4449, 7120, 5600, 3271, 8400, 8040, 4649, 3118, 1056, 5520, 4025, 2968, 0704, 8096, 3104, 2821, 7344, 6800, 2500, 2725, 8841, 7487, 1609, masses, american, age, although, argues, instead, states, dataset, enough, become, applicable, remain, valid, exception, occasionally, fraction, change, alter, results, appreciably, turns, cdot, represent, graphically, around, proceed, carefully, joint, band, hyperbolic, idea, likely, similarly, scriptstyle, gamma, sim, itself, proportionally, textstyle, latter, observations, sufficiently, classic, relies, either, allow, however, those, tell, precise, much, vary, specified, were, devised, plausible, repeated, very, times, 4pt, unknown, additionally, earlier, demonstrate, simplification, identity, simplified, 2x_, context, every, observation, say, mid, equiv, formalize, assertion, must, define, framework, corresponding, generated, plus, themselves, definition, requires, based, assuming, validity, evaluate, assumptions, discussed, needed, inhomogeneity, variances, extreme, self, evident, 401, uncorrelated, whether, meaning, goes, forced, framing, actually, type, issue, defines, our, had, expectation, positive, think, just, uniform, notice, constant, upfront, depend, deriving, showing, makes, connects, important, position, affect, connecting, multiplying, members, summation, numerator, thereby, details, concise, formula, generalizing, write, horizontal, indicate, shows, followup, expect, closer, phenomenon, toward, standardized, expressions, yields, reformulated, elements, 2ex, solved, directly, stand, alone, resultant, algebraically, ones, paragraph, below, proof, calculation, defining, respect, respective, introduced, derive, arguments, denoted, objective, solve, minimization, problem, find, provide, sense, mentioned, understood, minimizes, differences, actual, any, candidate, describes, hold, exactly, largely, suppose, observe, them, describe, common, stipulation, accuracy, corrected, concerns, conventionally, accurately, predicts, adjective, refers, outcome, predictor, cartesian, coordinate, gauss, markov, studentized, background, iteratively, reweighted, regularized, ridge, negative, local, multilevel, binary, choice, part, presumed, rate, macroeconomics, free, encyclopedia, projects, printable, download, print, export, switch, legacy, parser, shortened, url, cite, permanent, link, actions, english, українська, português, 한국어, 日本語, bahasa, indonesia, עברית, فارسی, eesti, ελληνικά, deutsch, català, беларуская, العربية, top, personal, special, pages, recent, community, learn, help, contribute, current, events, navigation, jump, content,
Text of the page (random words):
a frac sum _ i 1 n left x_ i bar x right left y_ i bar y right sum _ i 1 n left x_ i bar x right 2 frac sum _ i 1 n delta x_ i delta y_ i sum _ i 1 n delta x_ i 2 end aligned here we have introduced x displaystyle bar x and y displaystyle bar y as the average of the x i and y i respectively δ x i displaystyle delta x_ i and δ y i displaystyle delta y_ i as the deviations in x i and y i with respect to their respective means expanded formulas edit the above equations are efficient to use if the mean of the x and y variables x and y displaystyle bar x text and bar y are known if the means are not known at the time of calculation it may be more efficient to use the expanded version of the α and β displaystyle widehat alpha text and widehat beta equations these expanded equations may be derived from the more general polynomial regression equations 7 8 by defining the regression polynomial to be of order 1 as follows n i 1 n x i i 1 n x i i 1 n x i 2 α β i 1 n y i i 1 n y i x i displaystyle begin bmatrix n sum _ i 1 n x_ i 1ex sum _ i 1 n x_ i sum _ i 1 n x_ i 2 end bmatrix begin bmatrix widehat alpha 1ex widehat beta end bmatrix begin bmatrix sum _ i 1 n y_ i 1ex sum _ i 1 n y_ i x_ i end bmatrix the above system of linear equations may be solved directly or stand alone equations for α and β displaystyle widehat alpha text and widehat beta may be derived by expanding the matrix equations above the resultant equations are algebraically equivalent to the ones shown in the prior paragraph and are shown below without proof 9 7 α i 1 n y i i 1 n x i 2 i 1 n x i i 1 n x i y i n i 1 n x i 2 i 1 n x i 2 β n i 1 n x i y i i 1 n x i i 1 n y i n i 1 n x i 2 i 1 n x i 2 displaystyle begin aligned widehat alpha frac sum limits _ i 1 n y_ i sum limits _ i 1 n x_ i 2 sum limits _ i 1 n x_ i sum limits _ i 1 n x_ i y_ i n sum limits _ i 1 n x_ i 2 left sum limits _ i 1 n x_ i right 2 2ex widehat beta frac n sum limits _ i 1 n x_ i y_ i sum limits _ i 1 n x_ i sum limits _ i 1 n y_ i n sum limits _ i 1 n x_ i 2 left sum limits _ i 1 n x_ i right 2 end aligned interpretation edit relationship with the sample covariance matrix edit the solution can be reformulated using elements of the covariance matrix β s x y s x 2 r x y s y s x displaystyle widehat beta frac s_ x y s_ x 2 r_ xy frac s_ y s_ x where r xy is the sample correlation coefficient between x and y s x and s y are the uncorrected sample standard deviations of x and y s x 2 displaystyle s_ x 2 and s x y displaystyle s_ x y are the sample variance and sample covariance respectively substituting the above expressions for α displaystyle widehat alpha and β displaystyle widehat beta into the original solution yields y y s y r x y x x s x displaystyle frac y bar y s_ y r_ xy frac x bar x s_ x this shows that r xy is the slope of the regression line of the standardized data points and that this line passes through the origin since 1 r x y 1 displaystyle 1 leq r_ xy leq 1 then we get that if x is some measurement and y is a followup measurement from the same item then we expect that y on average will be closer to the mean measurement than it was to the original value of x this phenomenon is known as regressions toward the mean generalizing the x displaystyle bar x notation we can write a horizontal bar over an expression to indicate the average value of that expression over the set of samples for example x y 1 n i 1 n x i y i displaystyle overline xy frac 1 n sum _ i 1 n x_ i y_ i this notation allows us a concise formula for r xy r x y x y x y x 2 x 2 y 2 y 2 displaystyle r_ xy frac overline xy bar x bar y sqrt left overline x 2 bar x 2 right left overline y 2 bar y 2 right the coefficient of determination r squared is equal to r x y 2 displaystyle r_ xy 2 when the model is linear with a single independent variable see sample correlation coefficient for additional details interpretation about the slope edit by multiplying all members of the summation in the numerator by x i x x i x 1 displaystyle frac x_ i bar x x_ i bar x 1 thereby not changing it β i 1 n x i x y i y i 1 n x i x 2 i 1 n x i x 2 y i y x i x i 1 n x i x 2 i 1 n x i x 2 j 1 n x j x 2 y i y x i x displaystyle begin aligned widehat beta frac sum _ i 1 n left x_ i bar x right left y_ i bar y right sum _ i 1 n left x_ i bar x right 2 1ex frac sum _ i 1 n left x_ i bar x right 2 frac y_ i bar y x_ i bar x sum _ i 1 n left x_ i bar x right 2 1ex sum _ i 1 n frac left x_ i bar x right 2 sum _ j 1 n left x_ j bar x right 2 frac y_ i bar y x_ i bar x 6pt end aligned we can see that the slope tangent of angle of the regression line is the weighted average of y i y x i x displaystyle frac y_ i bar y x_ i bar x that is the slope tangent of angle of the line that connects the i th point to the average of all points weighted by x i x 2 displaystyle x_ i bar x 2 because the further the point is the more important it is since small errors in its position will affect the slope connecting it to the center point more interpretation about the intercept edit the parameter α displaystyle widehat alpha is the intercept of the linear function y α β x displaystyle begin aligned y widehat alpha widehat beta x 5pt end aligned therefore the y displaystyle y intercept of the function found with simple linear regression is y i n t e r c e p t α y β x displaystyle y_ rm intercept widehat alpha bar y widehat beta bar x because β displaystyle widehat beta is the slope of the linear function β tan θ displaystyle widehat beta tan theta therefore the angle θ displaystyle theta the graph of the function makes with the x displaystyle x axis is equal to θ arctan β displaystyle theta arctan widehat beta interpretation about the correlation edit in the above formulation notice that each x i displaystyle x_ i is a constant known upfront value while the y i displaystyle y_ i are random variables that depend on the linear function of x i displaystyle x_ i and the random term ε i displaystyle varepsilon _ i this assumption is used when deriving the standard error of the slope and showing that it is unbiased in this framing when x i displaystyle x_ i is not actually a random variable what type of parameter does the empirical correlation r x y displaystyle r_ xy estimate the issue is that for each value i we ll have e x i x i displaystyle e x_ i x_ i and v a r x i 0 displaystyle var x_ i 0 a possible interpretation of r x y displaystyle r_ xy is to imagine that x i displaystyle x_ i defines a random variable drawn from the empirical distribution of the x values in our sample for example if x had 10 values from the natural numbers 1 2 3 10 then we can imagine x to be a discrete uniform distribution under this interpretation all x i displaystyle x_ i have the same expectation and some positive variance with this interpretation we can think of r x y displaystyle r_ xy as the estimator of the pearson s correlation between the random variable y and the random variable x as we just defined it numerical properties edit the regression line goes through the center of mass point x y displaystyle bar x bar y if the model includes an intercept term i e not forced through the origin the sum of the residuals is zero if the model includes an intercept term i 1 n ε i 0 displaystyle sum _ i 1 n widehat varepsilon _ i 0 the residuals and x values are uncorrelated whether or not there is an intercept term in the model meaning i 1 n x i ε i 0 displaystyle sum _ i 1 n x_ i widehat varepsilon _ i 0 the relationship between ρ x y displaystyle rho _ xy the correlation coefficient for the population and the population variances of y displaystyle y σ y 2 displaystyle sigma _ y 2 and the error term of ε displaystyle varepsilon σ ε 2 displaystyle sigma _ varepsilon 2 is 10 401 σ ε 2 1 ρ x y 2 σ y 2 displaystyle sigma _ varepsilon 2 1 rho _ xy 2 sigma _ y 2 for extreme values of ρ x y displaystyle rho _ xy this is self evident since when ρ x y 0 displaystyle rho _ xy 0 then σ ε 2 σ y 2 displaystyle sigma _ varepsilon 2 sigma _ y 2 and when ρ x y 1 displaystyle rho _ xy 1 then σ ε 2 0 displaystyle sigma _ varepsilon 2 0 statistical properties edit description of the statistical properties of estimators from the simple linear regression estimates requires the use of a statistical model the following is based on assuming the validity of a model under which the estimates are optimal it is also possible to evaluate the properties under other assumptions such as inhomogeneity but this is discussed elsewhere clarification needed unbiasedness edit the estimators α displaystyle widehat alpha and β displaystyle widehat beta are unbiased to formalize this assertion we must define a framework in which these estimators are random variables we consider the residuals ε i as random variables drawn independently from some distribution with mean zero in other words for each value of x the corresponding value of y is generated as a mean response α βx plus an additional random variable ε called the error term equal to zero on average under such interpretation the least squares estimators α displaystyle widehat alpha and β displaystyle widehat beta will themselves be random variables whose means will equal the true values α and β this is the definition of an unbiased estimator variance of the mean response edit since the data in this context is defined to be x y pairs for every observation the mean response at a given value of x say x d is an estimate of the mean of the y values in the population at the x value of x d that is e y x d y d displaystyle hat e y mid x_ d equiv hat y _ d the variance of the mean response is given by 11 var α β x d var α var β x d 2 2 x d cov α β displaystyle operatorname var left hat alpha hat beta x_ d right operatorname var left hat alpha right left operatorname var hat beta right x_ d 2 2x_ d operatorname cov left hat alpha hat beta right this expression can be simplified to var α β x d σ 2 1 m x d x 2 x i x 2 displaystyle operatorname var left hat alpha hat beta x_ d right sigma 2 left frac 1 m frac left x_ d bar x right 2 sum x_ i bar x 2 right where m is the number of data points to demonstrate this simplification one can make use of the identity i x i x 2 i x i 2 1 m i x i 2 displaystyle sum _ i x_ i bar x 2 sum _ i x_ i 2 frac 1 m left sum _ i x_ i right 2 variance of the predicted response edit further information prediction interval the predicted response distribution is the predicted distribution of the residuals at the given point x d so the variance is given by var y d α β x d var y d var α β x d 2 cov y d α β x d var y d var α β x d displaystyle begin aligned operatorname var left y_ d left hat alpha hat beta x_ d right right operatorname var y_ d operatorname var left hat alpha hat beta x_ d right 2 operatorname cov left y_ d left hat alpha hat beta x_ d right right operatorname var y_ d operatorname var left hat alpha hat beta x_ d right end aligned the second line follows from the fact that cov y d α β x d displaystyle operatorname cov left y_ d left hat alpha hat beta x_ d right right is zero because the new prediction point is independent of the data used to fit the model additionally the term var α β x d displaystyle operatorname var left hat alpha hat beta x_ d right was calculated earlier for the mean response since var y d σ 2 displaystyle operatorname var y_ d sigma 2 a fixed but unknown parameter that can be estimated the variance of the predicted response is given by var y d α β x d σ 2 σ 2 1 m x d x 2 x i x 2 σ 2 1 1 m x d x 2 x i x 2 displaystyle begin aligned operatorname var left y_ d left hat alpha hat beta x_ d right right sigma 2 sigma 2 left frac 1 m frac left x_ d bar x right 2 sum x_ i bar x 2 right 4pt sigma 2 left 1 frac 1 m frac x_ d bar x 2 sum x_ i bar x 2 right end aligned confidence intervals edit the formulas given in the previous section allow one to calculate the point estimates of α and β that is the coefficients of the regression line for the given set of data however those formulas do not tell us how precise the estimates are i e how much the estimators α displaystyle widehat alpha and β displaystyle widehat beta vary from sample to sample for the specified sample size confidence intervals were devised to give a plausible set of values to the estimates one might have if one repeated the experiment a very large number of times the standard method of constructing confidence intervals for linear regression coefficients relies on the normality assumption which is justified if either the errors in the regression are normally distributed the so called classic regression assumption or the number of observations n is sufficiently large in which case the estimator is approximately normally distributed the latter case is justified by the central limit theorem normality assumption edit under the first assumption above that of the normality of the error terms the estimator of the slope coefficient will itself be normally distributed with mean β and variance σ 2 i x i x 2 textstyle sigma 2 left sum _ i x_ i bar x 2 right where σ 2 is the variance of the error terms see proofs involving ordinary least squares at the same time the sum of squared residuals q is distributed proportionally to χ 2 with n 2 degrees of freedom and independently from β displaystyle widehat beta this allows us to construct a t value t β β s β t n 2 displaystyle t frac widehat beta beta s_ widehat beta sim t_ n 2 where s β 1 n 2 i 1 n ε i 2 i 1 n x i x 2 displaystyle s_ widehat beta sqrt frac frac 1 n 2 sum _ i 1 n widehat varepsilon _ i 2 sum _ i 1 n x_ i bar x 2 is the unbiased standard error estimator of the estimator β displaystyle widehat beta this t value has a student s t distribution with n 2 degrees of freedom using it we can construct a confidence interval for β β β s β t n 2 β s β t n 2 displaystyle beta in left widehat beta s_ widehat beta t_ n 2 widehat beta s_ widehat beta t_ n 2 right at confidence level 1 γ where t n 2 displaystyle t_ n 2 is the 1 γ 2 th displaystyle scriptstyle left 1 frac gamma 2 right text th quantile of the t n 2 distribution for example if γ 0 05 then the confidence level is 95 similarly the confidence interval for the intercept coefficient α is given by α α s α t n 2 α s α t n 2 displaystyle alpha in left widehat alpha s_ widehat alpha t_ n 2 widehat alpha s_ widehat alpha t_ n 2 right at confidence level 1 γ where s α s β 1 n i 1 n x i 2 1 n n 2 i 1 n ε i 2 i 1 n x i 2 i 1 n x i x 2 displaystyle s_ widehat alpha s_ widehat beta sqrt frac 1 n sum _ i 1 n x_ i 2 sqrt frac 1 n n 2 left sum _ i 1 n widehat varepsilon _ i 2 right frac sum _ i 1 n x_ i 2 sum _ i 1 n x_ i bar x 2 the us changes in unemployment gdp growth regression with the 95 confidence bands the confidence intervals for α and β give us the general idea where these regression coefficients are most likely to be for example in the okun s law regression shown here the point estimates are α 0 859 β 1 817 displaystyle widehat alpha 0 859 qquad widehat beta 1 817 the 95 confidenc...
|