Skip to content
homethesisprojectswritingaboutworkquestionsarithmeticsequencesresume
Loading
~/blog/garch0%dark
  1. home/
  2. writing/
  3. Conditional Heteroskedasticity and GARCH

08 July 2026 · 5 min read · updated 23 Sept 2026

Conditional Heteroskedasticity and GARCH

A financial return is typically serially uncorrelated yet far from independent, the dependence living in its conditional variance rather than its conditional mean. Large moves cluster, so the squared return carries the autocorrelation the return itself lacks. We test for that effect, build the ARCH and GARCH recursions that model it, prove the heavy tails and the ARMA structure the squared process inherits, estimate the model by a one-pass conditional likelihood, and read persistence off the forecast. We close with the asymmetric leverage models, where a fall raises tomorrow's variance more than an equal rise.

  • 5 equations
  • 5 connections
  • time-series
  • econometrics
  • volatility
On this page▾
  • Detecting ARCH effects
  • The ARCH and GARCH models
  • Estimation
  • Persistence and forecasts
  • Asymmetry and the leverage effect

5 min left

  • Detecting ARCH effects1m
  • The ARCH and GARCH models1m
  • Estimation1m
  • Persistence and forecasts1m
  • Asymmetry and the leverage effect2m

A return is typically serially uncorrelated yet dependent, the dependence carried by its conditional variance rather than its conditional mean. Large moves cluster, so the squared return holds the autocorrelation the return itself lacks. Writing rt=μt+atr_t=\mu_t+a_trt​=μt​+at​ with at=σtεta_t=\sigma_t\varepsilon_tat​=σt​εt​, where {εt}\{\varepsilon_t\}{εt​} is iid with mean zero and unit variance and σt2=Var⁡(at∣Ft−1)\sigma_t^2=\Var(a_t\gvn\mathcal F_{t-1})σt2​=Var(at​∣Ft−1​) is the conditional variance, this post tests for that dependence and then models how the variance evolves [1], [2].

#Detecting ARCH effects

Volatility clustering shows up as autocorrelation in the squared returns even when the returns themselves are uncorrelated. The first task is to detect it, by testing whether past squared residuals help predict the current one.

::: theorem [thm

] In the auxiliary regression e^t2=α0+∑j=1qαje^t−j2+vt\hat e_t^2=\alpha_0+\sum_{j=1}^q\alpha_j\hat e_{t-j}^2+v_te^t2​=α0​+∑j=1q​αj​e^t−j2​+vt​, the score test of H0 ⁣:α1=⋯=αq=0H_0\!:\alpha_1=\dots=\alpha_q=0H0​:α1​=⋯=αq​=0 is LM=TR2→dχq2\mathrm{LM}=TR^2\cvD\chi^2_qLM=TR2d​χq2​. :::

::: proof The Gaussian score for the ARCH(q)(q)(q) parameters at H0H_0H0​ is proportional to the sample covariance of e^t2/σ^2−1\hat e_t^2/\hat\sigma^2-1e^t2​/σ^2−1 with the lagged e^t−j2\hat e_{t-j}^2e^t−j2​. The Lagrange-multiplier statistic equals TTT times the uncentred R2R^2R2 of the auxiliary regression, a quadratic form in qqq asymptotically Gaussian scores, hence χq2\chi^2_qχq2​. :::

ARCH-LM test for conditional heteroskedasticity
  input   residuals e_hat, lags q
  1  regress e_hat_t^2 on a constant and
     e_hat_{t-1}^2 .. e_hat_{t-q}^2
  2  LM <- T R^2 of the auxiliary regression
  3  reject the no-ARCH null when Prob(chi^2_q > LM) is small

#The ARCH and GARCH models

Detecting the effect is one thing, modelling its evolution another. The ARCH and GARCH families make the conditional variance an explicit function of past squared shocks and past variances.

::: definition [def

] The ARCH(m)(m)(m) model sets σt2=α0+∑i=1mαiat−i2\sigma_t^2=\alpha_0+\sum_{i=1}^m\alpha_i a_{t-i}^2σt2​=α0​+∑i=1m​αi​at−i2​ with α0>0\alpha_0>0α0​>0 and αi≥0\alpha_i\ge0αi​≥0. The GARCH(m,s)(m,s)(m,s) model adds lagged variances, σt2=α0+∑i=1mαiat−i2+∑j=1sβjσt−j2\sigma_t^2=\alpha_0+\sum_{i=1}^m\alpha_i a_{t-i}^2+\sum_{j=1}^s\beta_j\sigma_{t-j}^2σt2​=α0​+∑i=1m​αi​at−i2​+∑j=1s​βj​σt−j2​, with βj≥0\beta_j\ge0βj​≥0. :::

::: proposition [prop

] For the ARCH(1)(1)(1) model with 0≤α1<10\le\alpha_1<10≤α1​<1 the unconditional variance is Var⁡(at)=α0/(1−α1)\Var(a_t)=\alpha_0/(1-\alpha_1)Var(at​)=α0​/(1−α1​), and under Gaussian εt\varepsilon_tεt​ with 3α12<13\alpha_1^2<13α12​<1 the unconditional kurtosis is 3(1−α12)/(1−3α12)>33(1-\alpha_1^2)/(1-3\alpha_1^2)>33(1−α12​)/(1−3α12​)>3, so ata_tat​ is heavier tailed than its driving noise. :::

::: proof The mean is zero since E[at]=E[σt]E[εt]=0\E[a_t]=\E[\sigma_t]\E[\varepsilon_t]=0E[at​]=E[σt​]E[εt​]=0, the two factors being independent. By stationarity Var⁡(at)=E[σt2]=α0+α1E[at−12]=α0+α1Var⁡(at)\Var(a_t)=\E[\sigma_t^2]=\alpha_0+\alpha_1\E[a_{t-1}^2]=\alpha_0+\alpha_1\Var(a_t)Var(at​)=E[σt2​]=α0​+α1​E[at−12​]=α0​+α1​Var(at​), which gives Var⁡(at)=α0/(1−α1)\Var(a_t)=\alpha_0/(1-\alpha_1)Var(at​)=α0​/(1−α1​). Under normality E[at4∣Ft−1]=3σt4\E[a_t^4\gvn\mathcal F_{t-1}]=3\sigma_t^4E[at4​∣Ft−1​]=3σt4​, so m4=E[at4]=3E[(α0+α1at−12)2]m_4=\E[a_t^4]=3\E[(\alpha_0+\alpha_1 a_{t-1}^2)^2]m4​=E[at4​]=3E[(α0​+α1​at−12​)2]; solving the resulting stationary equation gives m4=3α02(1+α1)/[(1−α1)(1−3α12)]m_4=3\alpha_0^2(1+\alpha_1)/[(1-\alpha_1)(1-3\alpha_1^2)]m4​=3α02​(1+α1​)/[(1−α1​)(1−3α12​)], and m4/Var⁡(at)2m_4/\Var(a_t)^2m4​/Var(at​)2 is the stated kurtosis, which exceeds 333 for α1>0\alpha_1>0α1​>0. :::

The same accounting extends to GARCH once the recursion is rewritten in the squared series, where the lagged variances fold into an autoregressive term.

::: proposition [prop

] Let ηt=at2−σt2\eta_t=a_t^2-\sigma_t^2ηt​=at2​−σt2​. The GARCH(m,s)(m,s)(m,s) recursion is equivalent to

at2=α0+∑i=1max⁡(m,s)(αi+βi)at−i2+ηt−∑j=1sβjηt−j,(1)a_t^2=\alpha_0+\sum_{i=1}^{\max(m,s)}(\alpha_i+\beta_i)a_{t-i}^2+\eta_t-\sum_{j=1}^s\beta_j\eta_{t-j}, \tag*{(1)}at2​=α0​+i=1∑max(m,s)​(αi​+βi​)at−i2​+ηt​−j=1∑s​βj​ηt−j​,(1)

where {ηt}\{\eta_t\}{ηt​} is a mean-zero martingale difference, so the squared process is an ARMA. When ∑i(αi+βi)<1\sum_i(\alpha_i+\beta_i)<1∑i​(αi​+βi​)<1 it is covariance stationary with Var⁡(at)=α0/(1−∑i(αi+βi))\Var(a_t)=\alpha_0/\bigl(1-\sum_i(\alpha_i+\beta_i)\bigr)Var(at​)=α0​/(1−∑i​(αi​+βi​)). :::

::: proof Substitute σt−j2=at−j2−ηt−j\sigma_{t-j}^2=a_{t-j}^2-\eta_{t-j}σt−j2​=at−j2​−ηt−j​ into the GARCH recursion and collect the a2a^2a2 terms to reach Equation (1). The increment ηt=σt2(εt2−1)\eta_t=\sigma_t^2(\varepsilon_t^2-1)ηt​=σt2​(εt2​−1) has E[ηt∣Ft−1]=0\E[\eta_t\gvn\mathcal F_{t-1}]=0E[ηt​∣Ft−1​]=0, hence is a martingale difference and serially uncorrelated. Taking unconditional means of the ARMA equation, with E[ηt]=0\E[\eta_t]=0E[ηt​]=0 and stationarity, leaves E[at2]=α0+∑i(αi+βi)E[at2]\E[a_t^2]=\alpha_0+\sum_i(\alpha_i+\beta_i)\E[a_t^2]E[at2​]=α0​+∑i​(αi​+βi​)E[at2​], which rearranges to the stated variance. :::

The second moment settles whenever the persistence is below one, but the tails need the fourth moment, which exists only on a stricter region of the parameter space.

::: proposition [prop

] For the GARCH(1,1)(1,1)(1,1) model with iid εt\varepsilon_tεt​ of unit variance and finite fourth moment κ=E[εt4]\kappa=\E[\varepsilon_t^4]κ=E[εt4​], the fourth moment of ata_tat​ is finite if and only if (α1+β1)2+(κ−1)α12<1(\alpha_1+\beta_1)^2+(\kappa-1)\alpha_1^2<1(α1​+β1​)2+(κ−1)α12​<1, and then the kurtosis is

E[at4](E[at2])2=κ [1−(α1+β1)2]1−(α1+β1)2−(κ−1)α12,(2)\frac{\E[a_t^4]}{(\E[a_t^2])^2}=\frac{\kappa\,[1-(\alpha_1+\beta_1)^2]}{1-(\alpha_1+\beta_1)^2-(\kappa-1)\alpha_1^2}, \tag*{(2)}(E[at2​])2E[at4​]​=1−(α1​+β1​)2−(κ−1)α12​κ[1−(α1​+β1​)2]​,(2)

which exceeds κ\kappaκ whenever α1>0\alpha_1>0α1​>0, so the model adds tail weight beyond its innovation. :::

::: proof Write σt2=α0+ϕt−1σt−12\sigma_t^2=\alpha_0+\phi_{t-1}\sigma_{t-1}^2σt2​=α0​+ϕt−1​σt−12​ with ϕt−1=α1εt−12+β1\phi_{t-1}=\alpha_1\varepsilon_{t-1}^2+\beta_1ϕt−1​=α1​εt−12​+β1​, where ϕt−1\phi_{t-1}ϕt−1​ is a function of εt−1\varepsilon_{t-1}εt−1​ and so independent of σt−12∈Ft−2\sigma_{t-1}^2\in\mathcal F_{t-2}σt−12​∈Ft−2​. Then E[ϕ]=α1+β1\E[\phi]=\alpha_1+\beta_1E[ϕ]=α1​+β1​ and E[ϕ2]=κα12+2α1β1+β12=(α1+β1)2+(κ−1)α12\E[\phi^2]=\kappa\alpha_1^2+2\alpha_1\beta_1+\beta_1^2=(\alpha_1+\beta_1)^2+(\kappa-1)\alpha_1^2E[ϕ2]=κα12​+2α1​β1​+β12​=(α1​+β1​)2+(κ−1)α12​. Squaring the recursion and taking expectations under stationarity gives E[σt4]=α02+2α0E[ϕ]E[σt2]+E[ϕ2]E[σt4]\E[\sigma_t^4]=\alpha_0^2+2\alpha_0\E[\phi]\E[\sigma_t^2]+\E[\phi^2]\E[\sigma_t^4]E[σt4​]=α02​+2α0​E[ϕ]E[σt2​]+E[ϕ2]E[σt4​], which has a positive solution exactly when E[ϕ2]<1\E[\phi^2]<1E[ϕ2]<1. Substituting E[σt2]=α0/(1−α1−β1)\E[\sigma_t^2]=\alpha_0/(1-\alpha_1-\beta_1)E[σt2​]=α0​/(1−α1​−β1​) and E[at4]=κE[σt4]\E[a_t^4]=\kappa\E[\sigma_t^4]E[at4​]=κE[σt4​] and dividing by (E[at2])2(\E[a_t^2])^2(E[at2​])2 yields Equation (2), larger than κ\kappaκ once α1>0\alpha_1>0α1​>0. :::

#Estimation

Conditioning on the past turns the likelihood into a single forward pass, and the standardised residuals it leaves are what the diagnostics read.

::: proposition [prop

] Under εt∼N(0,1)\varepsilon_t\sim\Nor(0,1)εt​∼N(0,1) the parameter vector ϕ=(α0,…,βs)\phi=(\alpha_0,\dots,\beta_s)ϕ=(α0​,…,βs​) maximises the prediction-error decomposition

ℓ(ϕ)=−12∑t=1T(ln⁡σt2(ϕ)+at2σt2(ϕ)),(3)\ell(\phi)=-\tfrac12\sum_{t=1}^{T}\Bigl(\ln\sigma_t^2(\phi)+\frac{a_t^2}{\sigma_t^2(\phi)}\Bigr), \tag*{(3)}ℓ(ϕ)=−21​t=1∑T​(lnσt2​(ϕ)+σt2​(ϕ)at2​​),(3)

each σt2(ϕ)\sigma_t^2(\phi)σt2​(ϕ) computed by the model recursion from the past, so the likelihood is evaluated in a single forward pass. :::

::: proof The joint density factors as ∏tp(at∣Ft−1)\prod_t p(a_t\gvn\mathcal F_{t-1})∏t​p(at​∣Ft−1​) by the chain rule, and each conditional law is N(0,σt2)\Nor(0,\sigma_t^2)N(0,σt2​) with log density −12ln⁡(2π)−12ln⁡σt2−at2/(2σt2)-\tfrac12\ln(2\pi)-\tfrac12\ln\sigma_t^2-a_t^2/(2\sigma_t^2)−21​ln(2π)−21​lnσt2​−at2​/(2σt2​). Conditioning on Ft−1\mathcal F_{t-1}Ft−1​ makes σt2\sigma_t^2σt2​ measurable, so the recursion fixes it before ata_tat​ is read. Dropping the constant and summing the remaining terms gives Equation (3). :::

Maximum-likelihood estimation of GARCH(1,1)
  input   return shocks a_t, initial variance sigma_0^2
  1  for a candidate (alpha_0, alpha_1, beta_1) run the recursion
     sigma_t^2 = alpha_0 + alpha_1 a_{t-1}^2 + beta_1 sigma_{t-1}^2
  2  ell <- -1/2 sum_t ( ln sigma_t^2 + a_t^2 / sigma_t^2 )
  3  maximise ell over alpha_0>0, alpha_1,beta_1>=0, alpha_1+beta_1<1
  4  return the fit and standardised residuals a_t / sigma_t

::: remark Maximising the Gaussian ℓ\ellℓ stays consistent and asymptotically normal when the true innovation law is not Gaussian, the Gaussian score remaining an unbiased estimating equation for the variance dynamics; this is the quasi-maximum-likelihood estimator. A fitted model is checked through the standardised residuals ε^t=at/σ^t\hat\varepsilon_t=a_t/\hat\sigma_tε^t​=at​/σ^t​, which should be close to iid. The Ljung-Box statistic on ε^t\hat\varepsilon_tε^t​ probes the mean specification and on ε^t2\hat\varepsilon_t^2ε^t2​ the variance specification, ARCH effects surviving in ε^t2\hat\varepsilon_t^2ε^t2​ signalling an order set too low. :::

#Persistence and forecasts

How long a shock to variance lasts is set by the persistence of the recursion, which also fixes how a multistep variance forecast decays back toward its level.

::: proposition [prop

] For the GARCH(1,1)(1,1)(1,1) model the ℓ\ellℓ-step variance forecast satisfies σh2(ℓ)=α0+(α1+β1)σh2(ℓ−1)\sigma_h^2(\ell)=\alpha_0+(\alpha_1+\beta_1)\sigma_h^2(\ell-1)σh2​(ℓ)=α0​+(α1​+β1​)σh2​(ℓ−1) for ℓ>1\ell>1ℓ>1, so when α1+β1<1\alpha_1+\beta_1<1α1​+β1​<1 it converges to the unconditional variance α0/(1−α1−β1)\alpha_0/(1-\alpha_1-\beta_1)α0​/(1−α1​−β1​) as ℓ→∞\ell\to\inftyℓ→∞. When α1+β1=1\alpha_1+\beta_1=1α1​+β1​=1 it follows the line σh2(ℓ)=σh2(1)+(ℓ−1)α0\sigma_h^2(\ell)=\sigma_h^2(1)+(\ell-1)\alpha_0σh2​(ℓ)=σh2​(1)+(ℓ−1)α0​, the integrated GARCH case. :::

::: proof Write σt+12=α0+(α1+β1)σt2+α1σt2(εt2−1)\sigma_{t+1}^2=\alpha_0+(\alpha_1+\beta_1)\sigma_t^2+\alpha_1\sigma_t^2(\varepsilon_t^2-1)σt+12​=α0​+(α1​+β1​)σt2​+α1​σt2​(εt2​−1) using at2=σt2εt2a_t^2=\sigma_t^2\varepsilon_t^2at2​=σt2​εt2​. Since E[εt2−1∣Ft−1]=0\E[\varepsilon_t^2-1\gvn\mathcal F_{t-1}]=0E[εt2​−1∣Ft−1​]=0, the conditional expectation of the last term vanishes and the forecast obeys the affine recursion σh2(ℓ)=α0+(α1+β1)σh2(ℓ−1)\sigma_h^2(\ell)=\alpha_0+(\alpha_1+\beta_1)\sigma_h^2(\ell-1)σh2​(ℓ)=α0​+(α1​+β1​)σh2​(ℓ−1). A slope α1+β1<1\alpha_1+\beta_1<1α1​+β1​<1 contracts the map to its fixed point α0/(1−α1−β1)\alpha_0/(1-\alpha_1-\beta_1)α0​/(1−α1​−β1​), while at unit slope each step adds α0\alpha_0α0​, giving the line. :::

::: proposition [prop

] In the IGARCH(1,1)(1,1)(1,1) model α1+β1=1\alpha_1+\beta_1=1α1​+β1​=1 the recursion reads σt2=α0+(1−β1)at−12+β1σt−12\sigma_t^2=\alpha_0+(1-\beta_1)a_{t-1}^2+\beta_1\sigma_{t-1}^2σt2​=α0​+(1−β1​)at−12​+β1​σt−12​, and with α0=0\alpha_0=0α0​=0 it is the exponentially weighted moving average σt2=(1−β1)∑i≥0β1iat−1−i2\sigma_t^2=(1-\beta_1)\sum_{i\ge0}\beta_1^i a_{t-1-i}^2σt2​=(1−β1​)∑i≥0​β1i​at−1−i2​, whose multistep forecast is flat, σh2(ℓ)=σh2(1)\sigma_h^2(\ell)=\sigma_h^2(1)σh2​(ℓ)=σh2​(1) for every ℓ\ellℓ. :::

::: proof Setting α1=1−β1\alpha_1=1-\beta_1α1​=1−β1​ in the GARCH(1,1)(1,1)(1,1) recursion gives the displayed form. With α0=0\alpha_0=0α0​=0, unrolling σt2=(1−β1)at−12+β1σt−12\sigma_t^2=(1-\beta_1)a_{t-1}^2+\beta_1\sigma_{t-1}^2σt2​=(1−β1​)at−12​+β1​σt−12​ backward produces geometric weights (1−β1)β1i(1-\beta_1)\beta_1^i(1−β1​)β1i​ that sum to one. The forecast recursion of ?? at unit slope with α0=0\alpha_0=0α0​=0 holds every horizon at σh2(1)\sigma_h^2(1)σh2​(1). :::

#Asymmetry and the leverage effect

The GARCH variance enters past shocks only through their square, so it responds identically to a gain and a loss of equal size. Equity-like data instead carry a leverage effect, a larger variance after a drop, and an asymmetric law is needed to capture it.

::: definition [def

] The news impact curve is the function carrying the most recent shock at−1a_{t-1}at−1​ to σt2\sigma_t^2σt2​ with all earlier information held at its mean. For GARCH it is the parabola σt2=const+α1at−12\sigma_t^2=\text{const}+\alpha_1 a_{t-1}^2σt2​=const+α1​at−12​, symmetric about the origin. :::

::: definition [def

] The threshold, or GJR, model adds a half-on quadratic term,

σt2=α0+(α1+γ 1{at−1<0})at−12+β1σt−12,γ≥0,(4)\sigma_t^2=\alpha_0+(\alpha_1+\gamma\,\one\{a_{t-1}<0\})a_{t-1}^2+\beta_1\sigma_{t-1}^2,\qquad\gamma\ge0, \tag*{(4)}σt2​=α0​+(α1​+γ1{at−1​<0})at−12​+β1​σt−12​,γ≥0,(4)

so a negative shock loads on α1+γ\alpha_1+\gammaα1​+γ and a positive one on α1\alpha_1α1​ [3]. :::

::: proposition [prop

] With innovations symmetric about zero the threshold model is covariance stationary if and only if α1+γ/2+β1<1\alpha_1+\gamma/2+\beta_1<1α1​+γ/2+β1​<1, and then Var⁡(at)=α0/(1−α1−γ/2−β1)\Var(a_t)=\alpha_0/(1-\alpha_1-\gamma/2-\beta_1)Var(at​)=α0​/(1−α1​−γ/2−β1​). :::

::: proof Symmetry of εt\varepsilon_tεt​ gives E[1{at−1<0}at−12]=12E[at−12]\E[\one\{a_{t-1}<0\}a_{t-1}^2]=\tfrac12\E[a_{t-1}^2]E[1{at−1​<0}at−12​]=21​E[at−12​], the sign of at−1a_{t-1}at−1​ being that of εt−1\varepsilon_{t-1}εt−1​, independent of σt−1\sigma_{t-1}σt−1​ and equally likely either way. Taking unconditional expectations of the recursion under stationarity gives Var⁡(at)=α0+(α1+γ/2+β1)Var⁡(at)\Var(a_t)=\alpha_0+(\alpha_1+\gamma/2+\beta_1)\Var(a_t)Var(at​)=α0​+(α1​+γ/2+β1​)Var(at​), which rearranges to the stated variance and is positive exactly when the coefficient sum is below one. :::

::: proposition [prop

] In the threshold model with symmetric innovations and γ>0\gamma>0γ>0, Cov⁡(at−1,σt2)=γ E[at−131{at−1<0}]<0\operatorname{Cov}(a_{t-1},\sigma_t^2)=\gamma\,\E[a_{t-1}^3\one\{a_{t-1}<0\}]<0Cov(at−1​,σt2​)=γE[at−13​1{at−1​<0}]<0, so a fall raises next period's variance more than an equal rise. :::

::: proof Only the two shock terms of σt2\sigma_t^2σt2​ correlate with at−1a_{t-1}at−1​. Symmetry makes E[at−13]=0\E[a_{t-1}^3]=0E[at−13​]=0, so Cov⁡(at−1,α1at−12)=0\operatorname{Cov}(a_{t-1},\alpha_1 a_{t-1}^2)=0Cov(at−1​,α1​at−12​)=0, leaving Cov⁡(at−1,σt2)=γE[at−131{at−1<0}]\operatorname{Cov}(a_{t-1},\sigma_t^2)=\gamma\E[a_{t-1}^3\one\{a_{t-1}<0\}]Cov(at−1​,σt2​)=γE[at−13​1{at−1​<0}]. The integrand is negative wherever it is nonzero, so for a nondegenerate shock with finite third moment the expectation is strictly negative when γ>0\gamma>0γ>0. :::

A second route to asymmetry models the log variance, which also frees the sign of the coefficients.

::: definition [def

] The exponential GARCH models the log variance as an autoregression driven by g(εt−1)=θεt−1+γ(∣εt−1∣−E∣εt−1∣)g(\varepsilon_{t-1})=\theta\varepsilon_{t-1}+\gamma(|\varepsilon_{t-1}|-\E|\varepsilon_{t-1}|)g(εt−1​)=θεt−1​+γ(∣εt−1​∣−E∣εt−1​∣),

ln⁡σt2=α0+βln⁡σt−12+g(εt−1),(5)\ln\sigma_t^2=\alpha_0+\beta\ln\sigma_{t-1}^2+g(\varepsilon_{t-1}), \tag*{(5)}lnσt2​=α0​+βlnσt−12​+g(εt−1​),(5)

whose slope in εt−1\varepsilon_{t-1}εt−1​ is θ+γ\theta+\gammaθ+γ after a rise and θ−γ\theta-\gammaθ−γ after a fall [4]. :::

::: proposition [prop

] The forcing g(εt−1)g(\varepsilon_{t-1})g(εt−1​) is iid with mean zero, so ln⁡σt2\ln\sigma_t^2lnσt2​ is a stationary autoregression whenever ∣β∣<1|\beta|<1∣β∣<1, with E[ln⁡σt2]=α0/(1−β)\E[\ln\sigma_t^2]=\alpha_0/(1-\beta)E[lnσt2​]=α0​/(1−β), and modelling the log places no sign constraint on the parameters. :::

::: proof E[g(ε)]=θE[ε]+γ(E∣ε∣−E∣ε∣)=0\E[g(\varepsilon)]=\theta\E[\varepsilon]+\gamma(\E|\varepsilon|-\E|\varepsilon|)=0E[g(ε)]=θE[ε]+γ(E∣ε∣−E∣ε∣)=0, and g(εt−1)g(\varepsilon_{t-1})g(εt−1​) inherits the independence of εt−1\varepsilon_{t-1}εt−1​. The recursion is then an autoregression with iid mean-zero forcing, stationary for ∣β∣<1|\beta|<1∣β∣<1 with mean α0/(1−β)\alpha_0/(1-\beta)α0​/(1−β). Since σt2=exp⁡(ln⁡σt2)>0\sigma_t^2=\exp(\ln\sigma_t^2)>0σt2​=exp(lnσt2​)>0 for any real right-hand side, no nonnegativity constraint is needed. :::

GARCH is thus the variance-dynamics engine of financial time series. It turns the one fact a return's mean cannot show, that volatility clusters, into a parametric recursion whose persistence sets the half-life of a shock and whose asymmetric extensions carry the leverage effect. The clustering it models is exactly autocorrelation in the squared returns, and the conditional variance it produces is the discrete-time cousin of the stochastic-volatility diffusions that price options. A stochastic-volatility model, giving the log variance its own innovation rather than making it Ft−1\mathcal F_{t-1}Ft−1​-measurable, is the natural next step when a latent variance is wanted at the cost of a closed-form likelihood.

[1]
R. F. Engle, “Autoregressive conditional heteroscedasticity with estimates of the variance of United Kingdom inflation,” Econometrica, vol. 50, no. 4, pp. 987–1007, 1982.
[2]
T. Bollerslev, “Generalized autoregressive conditional heteroskedasticity,” Journal of Econometrics, vol. 31, no. 3, pp. 307–327, 1986.
[3]
L. R. Glosten, R. Jagannathan, and D. E. Runkle, “On the relation between the expected value and the volatility of the nominal excess return on stocks,” Journal of Finance, vol. 48, no. 5, pp. 1779–1801, 1993.
[4]
D. B. Nelson, “Conditional heteroskedasticity in asset returns: a new approach,” Econometrica, vol. 59, no. 2, pp. 347–370, 1991.

Part 6 of 6 in Econometrics

← previousCointegration

Explore connections

see in the atlas →

related

  • Serial Dependence: moving averages, ARMA, and long memory
  • Autocorrelation in Financial Time Series
  • Convergence and Limit Theorems

referenced by (1)

  • Serial Dependence: moving averages, ARMA, and long memory
cite
@misc{garch,
  author = {Zac Kienzle},
  title  = {Conditional Heteroskedasticity and GARCH},
  year   = {2026},
  month  = {07},
  url    = {https://zackienzle.com/blog/garch}
}