A financial return is typically serially uncorrelated yet far from independent,
the dependence living in its conditional variance rather than its conditional
mean. Large moves cluster, so the squared return carries the autocorrelation the
return itself lacks. We test for that effect, build the ARCH and GARCH
recursions that model it, prove the heavy tails and the ARMA structure the
squared process inherits, estimate the model by a one-pass conditional
likelihood, and read persistence off the forecast. We close with the asymmetric
leverage models, where a fall raises tomorrow's variance more than an equal
rise.
A return is typically serially uncorrelated yet dependent, the dependence carried by its conditional variance rather than its conditional mean. Large moves cluster, so the squared return holds the autocorrelation the return itself lacks. Writing rt=μt+at with at=σtεt, where {εt} is iid with mean zero and unit variance and σt2=Var(at∣Ft−1) is the conditional variance, this post tests for that dependence and then models how the variance evolves [1], [2].
Volatility clustering shows up as autocorrelation in the squared returns even when the returns themselves are uncorrelated. The first task is to detect it, by testing whether past squared residuals help predict the current one.
::: theorem [thm
]
In the auxiliary regression e^t2=α0+∑j=1qαje^t−j2+vt, the score test of H0:α1=⋯=αq=0 is LM=TR2dχq2.
:::
::: proof
The Gaussian score for the ARCH(q) parameters at H0 is proportional to the sample covariance of e^t2/σ^2−1 with the lagged e^t−j2. The Lagrange-multiplier statistic equals T times the uncentred R2 of the auxiliary regression, a quadratic form in q asymptotically Gaussian scores, hence χq2.
:::
ARCH-LM test for conditional heteroskedasticity
input residuals e_hat, lags q
1 regress e_hat_t^2 on a constant and
e_hat_{t-1}^2 .. e_hat_{t-q}^2
2 LM <- T R^2 of the auxiliary regression
3 reject the no-ARCH null when Prob(chi^2_q > LM) is small
Detecting the effect is one thing, modelling its evolution another. The ARCH and GARCH families make the conditional variance an explicit function of past squared shocks and past variances.
::: definition [def
]
The ARCH(m) model sets σt2=α0+∑i=1mαiat−i2 with α0>0 and αi≥0. The GARCH(m,s) model adds lagged variances, σt2=α0+∑i=1mαiat−i2+∑j=1sβjσt−j2, with βj≥0.
:::
::: proposition [prop
]
For the ARCH(1) model with 0≤α1<1 the unconditional variance is Var(at)=α0/(1−α1), and under Gaussian εt with 3α12<1 the unconditional kurtosis is 3(1−α12)/(1−3α12)>3, so at is heavier tailed than its driving noise.
:::
::: proof
The mean is zero since E[at]=E[σt]E[εt]=0, the two factors being independent. By stationarity Var(at)=E[σt2]=α0+α1E[at−12]=α0+α1Var(at), which gives Var(at)=α0/(1−α1). Under normality E[at4∣Ft−1]=3σt4, so m4=E[at4]=3E[(α0+α1at−12)2]; solving the resulting stationary equation gives m4=3α02(1+α1)/[(1−α1)(1−3α12)], and m4/Var(at)2 is the stated kurtosis, which exceeds 3 for α1>0.
:::
The same accounting extends to GARCH once the recursion is rewritten in the squared series, where the lagged variances fold into an autoregressive term.
::: proposition [prop
]
Let ηt=at2−σt2. The GARCH(m,s) recursion is equivalent to
where {ηt} is a mean-zero martingale difference, so the squared process is an ARMA. When ∑i(αi+βi)<1 it is covariance stationary with Var(at)=α0/(1−∑i(αi+βi)).
:::
::: proof
Substitute σt−j2=at−j2−ηt−j into the GARCH recursion and collect the a2 terms to reach Equation (1). The increment ηt=σt2(εt2−1) has E[ηt∣Ft−1]=0, hence is a martingale difference and serially uncorrelated. Taking unconditional means of the ARMA equation, with E[ηt]=0 and stationarity, leaves E[at2]=α0+∑i(αi+βi)E[at2], which rearranges to the stated variance.
:::
The second moment settles whenever the persistence is below one, but the tails need the fourth moment, which exists only on a stricter region of the parameter space.
::: proposition [prop
]
For the GARCH(1,1) model with iid εt of unit variance and finite fourth moment κ=E[εt4], the fourth moment of at is finite if and only if (α1+β1)2+(κ−1)α12<1, and then the kurtosis is
which exceeds κ whenever α1>0, so the model adds tail weight beyond its innovation.
:::
::: proof
Write σt2=α0+ϕt−1σt−12 with ϕt−1=α1εt−12+β1, where ϕt−1 is a function of εt−1 and so independent of σt−12∈Ft−2. Then E[ϕ]=α1+β1 and E[ϕ2]=κα12+2α1β1+β12=(α1+β1)2+(κ−1)α12. Squaring the recursion and taking expectations under stationarity gives E[σt4]=α02+2α0E[ϕ]E[σt2]+E[ϕ2]E[σt4], which has a positive solution exactly when E[ϕ2]<1. Substituting E[σt2]=α0/(1−α1−β1) and E[at4]=κE[σt4] and dividing by (E[at2])2 yields Equation (2), larger than κ once α1>0.
:::
Conditioning on the past turns the likelihood into a single forward pass, and the standardised residuals it leaves are what the diagnostics read.
::: proposition [prop
]
Under εt∼N(0,1) the parameter vector ϕ=(α0,…,βs) maximises the prediction-error decomposition
ℓ(ϕ)=−21t=1∑T(lnσt2(ϕ)+σt2(ϕ)at2),(3)
each σt2(ϕ) computed by the model recursion from the past, so the likelihood is evaluated in a single forward pass.
:::
::: proof
The joint density factors as ∏tp(at∣Ft−1) by the chain rule, and each conditional law is N(0,σt2) with log density −21ln(2π)−21lnσt2−at2/(2σt2). Conditioning on Ft−1 makes σt2 measurable, so the recursion fixes it before at is read. Dropping the constant and summing the remaining terms gives Equation (3).
:::
Maximum-likelihood estimation of GARCH(1,1)
input return shocks a_t, initial variance sigma_0^2
1 for a candidate (alpha_0, alpha_1, beta_1) run the recursion
sigma_t^2 = alpha_0 + alpha_1 a_{t-1}^2 + beta_1 sigma_{t-1}^2
2 ell <- -1/2 sum_t ( ln sigma_t^2 + a_t^2 / sigma_t^2 )
3 maximise ell over alpha_0>0, alpha_1,beta_1>=0, alpha_1+beta_1<1
4 return the fit and standardised residuals a_t / sigma_t
::: remark
Maximising the Gaussian ℓ stays consistent and asymptotically normal when the true innovation law is not Gaussian, the Gaussian score remaining an unbiased estimating equation for the variance dynamics; this is the quasi-maximum-likelihood estimator. A fitted model is checked through the standardised residuals ε^t=at/σ^t, which should be close to iid. The Ljung-Box statistic on ε^t probes the mean specification and on ε^t2 the variance specification, ARCH effects surviving in ε^t2 signalling an order set too low.
:::
How long a shock to variance lasts is set by the persistence of the recursion, which also fixes how a multistep variance forecast decays back toward its level.
::: proposition [prop
]
For the GARCH(1,1) model the ℓ-step variance forecast satisfies σh2(ℓ)=α0+(α1+β1)σh2(ℓ−1) for ℓ>1, so when α1+β1<1 it converges to the unconditional variance α0/(1−α1−β1) as ℓ→∞. When α1+β1=1 it follows the line σh2(ℓ)=σh2(1)+(ℓ−1)α0, the integrated GARCH case.
:::
::: proof
Write σt+12=α0+(α1+β1)σt2+α1σt2(εt2−1) using at2=σt2εt2. Since E[εt2−1∣Ft−1]=0, the conditional expectation of the last term vanishes and the forecast obeys the affine recursion σh2(ℓ)=α0+(α1+β1)σh2(ℓ−1). A slope α1+β1<1 contracts the map to its fixed point α0/(1−α1−β1), while at unit slope each step adds α0, giving the line.
:::
::: proposition [prop
]
In the IGARCH(1,1) model α1+β1=1 the recursion reads σt2=α0+(1−β1)at−12+β1σt−12, and with α0=0 it is the exponentially weighted moving average σt2=(1−β1)∑i≥0β1iat−1−i2, whose multistep forecast is flat, σh2(ℓ)=σh2(1) for every ℓ.
:::
::: proof
Setting α1=1−β1 in the GARCH(1,1) recursion gives the displayed form. With α0=0, unrolling σt2=(1−β1)at−12+β1σt−12 backward produces geometric weights (1−β1)β1i that sum to one. The forecast recursion of ?? at unit slope with α0=0 holds every horizon at σh2(1).
:::
The GARCH variance enters past shocks only through their square, so it responds identically to a gain and a loss of equal size. Equity-like data instead carry a leverage effect, a larger variance after a drop, and an asymmetric law is needed to capture it.
::: definition [def
]
The news impact curve is the function carrying the most recent shock at−1 to σt2 with all earlier information held at its mean. For GARCH it is the parabola σt2=const+α1at−12, symmetric about the origin.
:::
::: definition [def
]
The threshold, or GJR, model adds a half-on quadratic term,
so a negative shock loads on α1+γ and a positive one on α1[3].
:::
::: proposition [prop
]
With innovations symmetric about zero the threshold model is covariance stationary if and only if α1+γ/2+β1<1, and then Var(at)=α0/(1−α1−γ/2−β1).
:::
::: proof
Symmetry of εt gives E[1{at−1<0}at−12]=21E[at−12], the sign of at−1 being that of εt−1, independent of σt−1 and equally likely either way. Taking unconditional expectations of the recursion under stationarity gives Var(at)=α0+(α1+γ/2+β1)Var(at), which rearranges to the stated variance and is positive exactly when the coefficient sum is below one.
:::
::: proposition [prop
]
In the threshold model with symmetric innovations and γ>0, Cov(at−1,σt2)=γE[at−131{at−1<0}]<0, so a fall raises next period's variance more than an equal rise.
:::
::: proof
Only the two shock terms of σt2 correlate with at−1. Symmetry makes E[at−13]=0, so Cov(at−1,α1at−12)=0, leaving Cov(at−1,σt2)=γE[at−131{at−1<0}]. The integrand is negative wherever it is nonzero, so for a nondegenerate shock with finite third moment the expectation is strictly negative when γ>0.
:::
A second route to asymmetry models the log variance, which also frees the sign of the coefficients.
::: definition [def
]
The exponential GARCH models the log variance as an autoregression driven by g(εt−1)=θεt−1+γ(∣εt−1∣−E∣εt−1∣),
lnσt2=α0+βlnσt−12+g(εt−1),(5)
whose slope in εt−1 is θ+γ after a rise and θ−γ after a fall [4].
:::
::: proposition [prop
]
The forcing g(εt−1) is iid with mean zero, so lnσt2 is a stationary autoregression whenever ∣β∣<1, with E[lnσt2]=α0/(1−β), and modelling the log places no sign constraint on the parameters.
:::
::: proof
E[g(ε)]=θE[ε]+γ(E∣ε∣−E∣ε∣)=0, and g(εt−1) inherits the independence of εt−1. The recursion is then an autoregression with iid mean-zero forcing, stationary for ∣β∣<1 with mean α0/(1−β). Since σt2=exp(lnσt2)>0 for any real right-hand side, no nonnegativity constraint is needed.
:::
GARCH is thus the variance-dynamics engine of financial time series. It turns the one fact a return's mean cannot show, that volatility clusters, into a parametric recursion whose persistence sets the half-life of a shock and whose asymmetric extensions carry the leverage effect. The clustering it models is exactly autocorrelation in the squared returns, and the conditional variance it produces is the discrete-time cousin of the stochastic-volatility diffusions that price options. A stochastic-volatility model, giving the log variance its own innovation rather than making it Ft−1-measurable, is the natural next step when a latent variance is wanted at the cost of a closed-form likelihood.
[1]
R. F. Engle, “Autoregressive conditional heteroscedasticity with estimates of the variance of United Kingdom inflation,” Econometrica, vol. 50, no. 4, pp. 987–1007, 1982.
[2]
T. Bollerslev, “Generalized autoregressive conditional heteroskedasticity,” Journal of Econometrics, vol. 31, no. 3, pp. 307–327, 1986.
[3]
L. R. Glosten, R. Jagannathan, and D. E. Runkle, “On the relation between the expected value and the volatility of the nominal excess return on stocks,” Journal of Finance, vol. 48, no. 5, pp. 1779–1801, 1993.
[4]
D. B. Nelson, “Conditional heteroskedasticity in asset returns: a new approach,” Econometrica, vol. 59, no. 2, pp. 347–370, 1991.