Skip to content
homethesisprojectswritingaboutworkquestionsarithmeticsequencesresume
Loading
~/blog/serial-dependence0%dark
  1. home/
  2. writing/
  3. Serial Dependence: moving averages, ARMA, and long memory

08 July 2026 · 6 min read · updated 23 Sept 2026

Serial Dependence: moving averages, ARMA, and long memory

The autoregression is one half of linear time-series analysis. Its dual is the moving average, whose autocorrelation truncates rather than decays, and their synthesis is the ARMA model, whose parsimony comes from a ratio of polynomials. We build the moving average and its invertibility, derive the two long-division representations of an ARMA model and the impulse responses they expose, read the seasonal airline model off its autocorrelations, and push dependence past the geometric rate into fractionally integrated long memory with hyperbolic decay. We close with the regression whose errors are serially correlated and the heteroskedasticity- and autocorrelation-consistent covariance that keeps its inference valid.

  • 5 equations
  • 5 connections
  • time-series
  • econometrics
  • stochastic-processes
On this page▾
  • Moving averages
  • ARMA models
  • Three representations
  • Seasonality
  • Long memory
  • Inference under serial correlation

6 min left

  • Moving averages1m
  • ARMA models1m
  • Three representations1m
  • Seasonality1m
  • Long memory1m
  • Inference under serial correlation1m

The autoregression captures dependence that decays. Its dual is the moving average, whose dependence truncates, and their synthesis is the ARMA model, whose economy comes from a ratio of two polynomials. This post completes the linear picture the autocorrelation and autoregression posts began. It builds the moving average and its invertibility, the ARMA model and the two representations long division exposes, the seasonal airline model, and the fractionally integrated processes whose autocorrelation decays only hyperbolically, then closes with valid inference for a regression whose errors carry serial correlation [1], [2].

#Moving averages

An autoregression is an infinite moving average with geometrically constrained weights. The moving average takes the opposite tack, a finite sum of shocks.

::: definition [def

] The MA(qqq) model is rt=c0+θ(B)atr_t = c_0 + \theta(B)a_trt​=c0​+θ(B)at​ with θ(B)=1−θ1B−⋯−θqBq\theta(B) = 1 - \theta_1 B - \cdots - \theta_q B^qθ(B)=1−θ1​B−⋯−θq​Bq, backshift BBB, and {at}\{a_t\}{at​} white noise of variance σ2\sigma^2σ2. It is a finite linear combination of white noise, so it is weakly stationary for every parameter value. :::

::: proposition [prop

] The MA(qqq) model has mean c0c_0c0​, variance (1+θ12+⋯+θq2)σ2(1 + \theta_1^2 + \cdots + \theta_q^2)\sigma^2(1+θ12​+⋯+θq2​)σ2, and autocorrelations that vanish beyond lag qqq, so its autocorrelation function cuts off at qqq. :::

::: proof Taking expectations gives E[rt]=c0\E[r_t] = c_0E[rt​]=c0​ since E[at−i]=0\E[a_{t-i}] = 0E[at−i​]=0. Centre and use that the at−ia_{t-i}at−i​ are uncorrelated with common variance σ2\sigma^2σ2, so Var⁡(rt)=σ2∑i=0qθi2\Var(r_t) = \sigma^2\sum_{i=0}^q \theta_i^2Var(rt​)=σ2∑i=0q​θi2​ with θ0=1\theta_0 = 1θ0​=1. For ℓ>q\ell > qℓ>q the windows {at,…,at−q}\{a_t,\dots,a_{t-q}\}{at​,…,at−q​} and {at−ℓ,…,at−ℓ−q}\{a_{t-\ell},\dots,a_{t-\ell-q}\}{at−ℓ​,…,at−ℓ−q​} are disjoint, so every cross term E[asau]\E[a_s a_u]E[as​au​] with s≠us\ne us=u vanishes and γℓ=0\gamma_\ell = 0γℓ​=0. For 1≤ℓ≤q1 \le \ell \le q1≤ℓ≤q the overlap leaves a finite combination of the θj\theta_jθj​, and dividing by γ0\gamma_0γ0​ gives an autocorrelation that cuts off at qqq. :::

The cutoff is the mirror of the autoregression, whose partial autocorrelation cuts off while its autocorrelation decays. The moving average raises a dual question, whether the shocks can be recovered from the observed series.

::: lemma [lem

] The MA(qqq) model rt=θ(B)atr_t = \theta(B)a_trt​=θ(B)at​ admits the autoregressive form π(B)rt=at\pi(B)r_t = a_tπ(B)rt​=at​ with summable weights π\piπ if and only if every root of θ(z)=0\theta(z) = 0θ(z)=0 lies outside the unit circle. That condition makes the shock ata_tat​ recoverable from the present and past of rtr_trt​, and the model is then called invertible. :::

::: proof Formally π(B)=θ(B)−1=∏i(1−ηiB)−1\pi(B) = \theta(B)^{-1} = \prod_i (1 - \eta_i B)^{-1}π(B)=θ(B)−1=∏i​(1−ηi​B)−1 where the ηi\eta_iηi​ are the reciprocals of the roots of θ\thetaθ. Each factor expands as ∑k≥0ηikBk\sum_{k\ge0}\eta_i^k B^k∑k≥0​ηik​Bk, which converges with summable coefficients exactly when ∣ηi∣<1|\eta_i| < 1∣ηi​∣<1, that is when every root 1/ηi1/\eta_i1/ηi​ lies outside the unit circle. The resulting π(B)rt=at\pi(B)r_t = a_tπ(B)rt​=at​ then writes ata_tat​ as a convergent combination of current and past rtr_trt​ [1]. :::

#ARMA models

Combining the two mechanisms economises parameters. Where a pure AR or MA might need a high order, their ratio can be short.

::: definition [def

] The ARMA(p,qp,qp,q) model is ϕ(B)rt=ϕ0+θ(B)at\phi(B)r_t = \phi_0 + \theta(B)a_tϕ(B)rt​=ϕ0​+θ(B)at​ with ϕ(B)=1−ϕ1B−⋯−ϕpBp\phi(B) = 1 - \phi_1 B - \cdots - \phi_p B^pϕ(B)=1−ϕ1​B−⋯−ϕp​Bp and θ(B)=1−θ1B−⋯−θqBq\theta(B) = 1 - \theta_1 B - \cdots - \theta_q B^qθ(B)=1−θ1​B−⋯−θq​Bq sharing no common factor. It is weakly stationary if and only if every root of ϕ(z)=0\phi(z) = 0ϕ(z)=0 lies outside the unit circle, with mean μ=ϕ0/ϕ(1)\mu = \phi_0/\phi(1)μ=ϕ0​/ϕ(1). :::

The mixed model's autocorrelation behaves like an autoregression after a short transient set by the moving-average part.

::: proposition [prop

] The stationary ARMA(1,11,11,1) model rt=ϕ1rt−1+ϕ0+at−θ1at−1r_t = \phi_1 r_{t-1} + \phi_0 + a_t - \theta_1 a_{t-1}rt​=ϕ1​rt−1​+ϕ0​+at​−θ1​at−1​ with ∣ϕ1∣<1|\phi_1| < 1∣ϕ1​∣<1 and ϕ1≠θ1\phi_1 \ne \theta_1ϕ1​=θ1​ has

γ0=1−2ϕ1θ1+θ121−ϕ12 σ2,γ1=ϕ1γ0−θ1σ2,γℓ=ϕ1γℓ−1 (ℓ≥2).(1)\gamma_0 = \frac{1 - 2\phi_1\theta_1 + \theta_1^2}{1 - \phi_1^2}\,\sigma^2,\qquad \gamma_1 = \phi_1\gamma_0 - \theta_1\sigma^2,\qquad \gamma_\ell = \phi_1\gamma_{\ell-1}\ (\ell \ge 2). \tag*{(1)}γ0​=1−ϕ12​1−2ϕ1​θ1​+θ12​​σ2,γ1​=ϕ1​γ0​−θ1​σ2,γℓ​=ϕ1​γℓ−1​ (ℓ≥2).(1)

Its autocorrelation therefore decays geometrically at rate ϕ1\phi_1ϕ1​ but only from lag 2 onward, so neither the autocorrelation nor the partial autocorrelation cuts off at any finite lag. :::

::: proof Since rt−1r_{t-1}rt−1​ and ata_tat​ are uncorrelated, E[rtat]=E[at2]=σ2\E[r_t a_t] = \E[a_t^2] = \sigma^2E[rt​at​]=E[at2​]=σ2 and E[rtat−1]=ϕ1σ2−θ1σ2\E[r_t a_{t-1}] = \phi_1\sigma^2 - \theta_1\sigma^2E[rt​at−1​]=ϕ1​σ2−θ1​σ2. Taking the variance of the model, γ0=ϕ12γ0+(1+θ12)σ2+2ϕ1Cov⁡(rt−1,at−θ1at−1)=ϕ12γ0+(1+θ12)σ2−2ϕ1θ1σ2\gamma_0 = \phi_1^2\gamma_0 + (1+\theta_1^2)\sigma^2 + 2\phi_1\Cov(r_{t-1}, a_t - \theta_1 a_{t-1}) = \phi_1^2\gamma_0 + (1+\theta_1^2)\sigma^2 - 2\phi_1\theta_1\sigma^2γ0​=ϕ12​γ0​+(1+θ12​)σ2+2ϕ1​Cov(rt−1​,at​−θ1​at−1​)=ϕ12​γ0​+(1+θ12​)σ2−2ϕ1​θ1​σ2, using Cov⁡(rt−1,at−1)=σ2\Cov(r_{t-1}, a_{t-1}) = \sigma^2Cov(rt−1​,at−1​)=σ2. Rearranging gives γ0\gamma_0γ0​. Multiplying the model by rt−1r_{t-1}rt−1​ and taking expectations gives γ1=ϕ1γ0−θ1σ2\gamma_1 = \phi_1\gamma_0 - \theta_1\sigma^2γ1​=ϕ1​γ0​−θ1​σ2, the shock term surviving because E[at−1rt−1]=σ2\E[a_{t-1}r_{t-1}] = \sigma^2E[at−1​rt−1​]=σ2. Multiplying by rt−ℓr_{t-\ell}rt−ℓ​ for ℓ≥2\ell \ge 2ℓ≥2 annihilates both shock terms and leaves γℓ=ϕ1γℓ−1\gamma_\ell = \phi_1\gamma_{\ell-1}γℓ​=ϕ1​γℓ−1​, which is Equation (1). Because both the moving-average and the autoregressive weights are infinite once no common factor cancels, neither function truncates. :::

Since neither function cuts off, order identification cannot read ppp and qqq off a truncation as it does for the pure models. The extended autocorrelation function locates the corner of the (p,q)(p,q)(p,q) order in a two-way table by first estimating the autoregressive part consistently and reading the moving-average order off the derived residuals [3].

#Three representations

A stationary ARMA model has three equivalent forms, each serving a purpose, obtained by long division of the two polynomials.

::: proposition [prop

] Write ψ(B)=θ(B)/ϕ(B)\psi(B) = \theta(B)/\phi(B)ψ(B)=θ(B)/ϕ(B) and π(B)=ϕ(B)/θ(B)\pi(B) = \phi(B)/\theta(B)π(B)=ϕ(B)/θ(B). The ARMA(p,qp,qp,q) model has the moving-average representation rt=μ+ψ(B)atr_t = \mu + \psi(B)a_trt​=μ+ψ(B)at​, whose weights are the impulse responses and solve

ψ0=1,ψj=∑i=1pϕiψj−i−θj(j≥1),(2)\psi_0 = 1,\qquad \psi_j = \sum_{i=1}^p \phi_i\psi_{j-i} - \theta_j\quad(j\ge1), \tag*{(2)}ψ0​=1,ψj​=i=1∑p​ϕi​ψj−i​−θj​(j≥1),(2)

with θj=0\theta_j = 0θj​=0 for j>qj > qj>q, and the autoregressive representation π(B)rt=ϕ0/θ(1)+at\pi(B)r_t = \phi_0/\theta(1) + a_tπ(B)rt​=ϕ0​/θ(1)+at​ valid when the roots of θ\thetaθ lie outside the unit circle. Both weight sequences decay geometrically, so a remote shock and a remote observation each have vanishing influence. :::

::: proof From ϕ(B)ψ(B)=θ(B)\phi(B)\psi(B) = \theta(B)ϕ(B)ψ(B)=θ(B), matching the coefficient of BjB^jBj gives ψj−∑i=1pϕiψj−i=−θj\psi_j - \sum_{i=1}^p\phi_i\psi_{j-i} = -\theta_jψj​−∑i=1p​ϕi​ψj−i​=−θj​ with ψ0=1\psi_0 = 1ψ0​=1, which is Equation (2). The homogeneous part ψj=∑iϕiψj−i\psi_j = \sum_i\phi_i\psi_{j-i}ψj​=∑i​ϕi​ψj−i​ for j>qj > qj>q is the autoregressive recursion, so the ψj\psi_jψj​ inherit the geometric decay of the reciprocal roots of ϕ\phiϕ under stationarity. Dividing ϕ(B)\phi(B)ϕ(B) by θ(B)\theta(B)θ(B) instead yields π(B)\pi(B)π(B) symmetrically, with geometric decay of the πj\pi_jπj​ under invertibility, and Bc=cBc = cBc=c for a constant gives the intercept ϕ0/θ(1)\phi_0/\theta(1)ϕ0​/θ(1). For the ARMA(1,11,11,1) case the recursion solves in closed form, ψj=ϕ1 j−1(ϕ1−θ1)\psi_j = \phi_1^{\,j-1}(\phi_1 - \theta_1)ψj​=ϕ1j−1​(ϕ1​−θ1​) for j≥1j\ge1j≥1. :::

The moving-average representation gives the forecast error directly. The ℓ\ellℓ-step error is ∑j=0ℓ−1ψjah+ℓ−j\sum_{j=0}^{\ell-1}\psi_j a_{h+\ell-j}∑j=0ℓ−1​ψj​ah+ℓ−j​, so its variance is σ2∑j=0ℓ−1ψj2\sigma^2\sum_{j=0}^{\ell-1}\psi_j^2σ2∑j=0ℓ−1​ψj2​, rising to γ0=σ2∑j≥0ψj2\gamma_0 = \sigma^2\sum_{j\ge0}\psi_j^2γ0​=σ2∑j≥0​ψj2​, exactly the mean-reverting forecast the autoregression post derived. The two rational forms are the finite-parameter handle on a universal object.

::: remark [rem

] Every zero-mean weakly stationary purely nondeterministic series admits the Wold representation rt=∑j≥0ψjat−jr_t = \sum_{j\ge0}\psi_j a_{t-j}rt​=∑j≥0​ψj​at−j​ with ∑jψj2<∞\sum_j\psi_j^2 < \infty∑j​ψj2​<∞, proved in the autoregression post. The ARMA model is the rational case ψ(B)=θ(B)/ϕ(B)\psi(B) = \theta(B)/\phi(B)ψ(B)=θ(B)/ϕ(B), a finite-parameter approximation to this infinite-order representation [2]. :::

#Seasonality

A series sampled through a repeating calendar carries dependence at the seasonal lag alongside the short-run lag. The multiplicative model handles both with two moving-average factors.

::: definition [def

] For a series of period sss, the airline model is

(1−B)(1−Bs)xt=(1−θB)(1−ΘBs)at,∣θ∣<1, ∣Θ∣<1,(3)(1 - B)(1 - B^s)x_t = (1 - \theta B)(1 - \Theta B^s)a_t,\qquad |\theta| < 1,\ |\Theta| < 1, \tag*{(3)}(1−B)(1−Bs)xt​=(1−θB)(1−ΘBs)at​,∣θ∣<1, ∣Θ∣<1,(3)

a regular and a seasonal difference on the left and a regular times a seasonal moving average on the right. :::

::: proposition [prop

] The differenced series wt=(1−θB)(1−ΘBs)atw_t = (1 - \theta B)(1 - \Theta B^s)a_twt​=(1−θB)(1−ΘBs)at​ has autocorrelation nonzero only at lags 1, s−1s-1s−1, sss, and s+1s+1s+1, with

ρ1=−θ1+θ2,ρs=−Θ1+Θ2,ρs−1=ρs+1=ρ1ρs,(4)\rho_1 = \frac{-\theta}{1+\theta^2},\qquad \rho_s = \frac{-\Theta}{1+\Theta^2},\qquad \rho_{s-1} = \rho_{s+1} = \rho_1\rho_s, \tag*{(4)}ρ1​=1+θ2−θ​,ρs​=1+Θ2−Θ​,ρs−1​=ρs+1​=ρ1​ρs​,(4)

so the regular and seasonal dynamics enter approximately orthogonally. :::

::: proof Expand wt=at−θat−1−Θat−s+θΘat−s−1w_t = a_t - \theta a_{t-1} - \Theta a_{t-s} + \theta\Theta a_{t-s-1}wt​=at​−θat−1​−Θat−s​+θΘat−s−1​. Since the ata_tat​ are uncorrelated with variance σ2\sigma^2σ2, direct computation gives γ0=(1+θ2)(1+Θ2)σ2\gamma_0 = (1+\theta^2)(1+\Theta^2)\sigma^2γ0​=(1+θ2)(1+Θ2)σ2, γ1=−θ(1+Θ2)σ2\gamma_1 = -\theta(1+\Theta^2)\sigma^2γ1​=−θ(1+Θ2)σ2, γs=−Θ(1+θ2)σ2\gamma_s = -\Theta(1+\theta^2)\sigma^2γs​=−Θ(1+θ2)σ2, and γs−1=γs+1=θΘσ2\gamma_{s-1} = \gamma_{s+1} = \theta\Theta\sigma^2γs−1​=γs+1​=θΘσ2, with every other autocovariance zero because the shifted index windows share no innovation. Dividing by γ0\gamma_0γ0​ gives Equation (4), and the product form ρs−1=ρ1ρs\rho_{s-1} = \rho_1\rho_sρs−1​=ρ1​ρs​ is the signature of the multiplicative structure. Deterministic seasonality handled by dummy variables is the boundary case Θ=1\Theta = 1Θ=1 [1]. :::

#Long memory

Under stationarity the autocorrelation decays geometrically; under a unit root it does not decay at all. Between the two lies a process whose autocorrelation decays only at a polynomial rate.

::: definition [def

] The fractionally integrated process is (1−B)dxt=at(1 - B)^d x_t = a_t(1−B)dxt​=at​ with −12<d<12-\tfrac12 < d < \tfrac12−21​<d<21​, where the fractional difference expands by the binomial series (1−B)d=∑k≥0(dk)(−B)k(1 - B)^d = \sum_{k\ge0}\binom{d}{k}(-B)^k(1−B)d=∑k≥0​(kd​)(−B)k. If (1−B)dxt(1-B)^d x_t(1−B)dxt​ follows an ARMA(p,qp,qp,q) model, xtx_txt​ is an ARFIMA(p,d,qp,d,qp,d,q) process. :::

::: proposition [prop

] For 0<d<120 < d < \tfrac120<d<21​ the process is weakly stationary with moving-average weights ψk=Γ(k+d)/(Γ(k+1)Γ(d))∼kd−1/Γ(d)\psi_k = \Gamma(k+d)/(\Gamma(k+1)\Gamma(d)) \sim k^{d-1}/\Gamma(d)ψk​=Γ(k+d)/(Γ(k+1)Γ(d))∼kd−1/Γ(d) and autocorrelations ρk∼c k2d−1\rho_k \sim c\,k^{2d-1}ρk​∼ck2d−1, a hyperbolic decay slower than any geometric rate, while its spectral density diverges at the origin, f(ω)∼ω−2df(\omega) \sim \omega^{-2d}f(ω)∼ω−2d as ω→0\omega \to 0ω→0. :::

::: proof Expanding (1−B)−d=∑k(−dk)(−B)k(1-B)^{-d} = \sum_k\binom{-d}{k}(-B)^k(1−B)−d=∑k​(k−d​)(−B)k gives the stated ψk=Γ(k+d)/(Γ(k+1)Γ(d))\psi_k = \Gamma(k+d)/(\Gamma(k+1)\Gamma(d))ψk​=Γ(k+d)/(Γ(k+1)Γ(d)), and Stirling's ratio Γ(k+d)/Γ(k+1)∼kd−1\Gamma(k+d)/\Gamma(k+1) \sim k^{d-1}Γ(k+d)/Γ(k+1)∼kd−1 yields ψk∼kd−1/Γ(d)\psi_k \sim k^{d-1}/\Gamma(d)ψk​∼kd−1/Γ(d). These are square-summable exactly when 2(d−1)<−12(d-1) < -12(d−1)<−1, that is d<12d < \tfrac12d<21​, so the representation converges and the process is stationary. The autocovariance γk=σ2∑jψjψj+k\gamma_k = \sigma^2\sum_j\psi_j\psi_{j+k}γk​=σ2∑j​ψj​ψj+k​ is a tail sum of jd−1(j+k)d−1j^{d-1}(j+k)^{d-1}jd−1(j+k)d−1, of order k2d−1k^{2d-1}k2d−1 by comparison with the integral, giving the power law. For the spectrum, the transfer function of the pure fractional filter gives f(ω)=(σ2/2π)∣1−e−iω∣−2df(\omega) = (\sigma^2/2\pi)|1 - e^{-i\omega}|^{-2d}f(ω)=(σ2/2π)∣1−e−iω∣−2d, the frequency-domain complement of the AR spectrum in the autoregression post, and ∣1−e−iω∣2=4sin⁡2(ω/2)∼ω2|1 - e^{-i\omega}|^2 = 4\sin^2(\omega/2) \sim \omega^2∣1−e−iω∣2=4sin2(ω/2)∼ω2 as ω→0\omega \to 0ω→0 gives f(ω)∼(σ2/2π) ω−2d→∞f(\omega) \sim (\sigma^2/2\pi)\,\omega^{-2d} \to \inftyf(ω)∼(σ2/2π)ω−2d→∞ [4]. :::

The hyperbolic decay is not a curiosity. The sample autocorrelation of absolute or squared asset returns is small but positive over hundreds of lags, the empirical fingerprint of long memory in volatility [5].

#Inference under serial correlation

A regression whose errors carry serial correlation keeps a consistent coefficient but loses the textbook standard error. Consider yt=xt⊤β+ety_t = x_t^{\top}\beta + e_tyt​=xt⊤​β+et​ with ete_tet​ serially correlated. Ordinary least squares stays unbiased, yet its conventional covariance σ2(∑xtxt⊤)−1\sigma^2(\sum x_t x_t^{\top})^{-1}σ2(∑xt​xt⊤​)−1 is wrong, and the fix is a covariance robust to both heteroskedasticity and autocorrelation.

::: proposition [prop

] For ordinary least squares in yt=xt⊤β+ety_t = x_t^{\top}\beta + e_tyt​=xt⊤​β+et​, T(β^−β)→N(0,Q−1SQ−1)\sqrt{T}(\hat\beta - \beta) \to \mathcal N(0, Q^{-1} S Q^{-1})T​(β^​−β)→N(0,Q−1SQ−1) with Q=E[xtxt⊤]Q = \E[x_t x_t^{\top}]Q=E[xt​xt⊤​] and SSS the long-run variance of xtetx_t e_txt​et​. The Newey-West estimator

S^=Γ^0+∑ℓ=1Lwℓ(Γ^ℓ+Γ^ℓ⊤),wℓ=1−ℓL+1,(5)\hat S = \hat\Gamma_0 + \sum_{\ell=1}^{L} w_\ell\bigl(\hat\Gamma_\ell + \hat\Gamma_\ell^{\top}\bigr),\qquad w_\ell = 1 - \frac{\ell}{L+1}, \tag*{(5)}S^=Γ^0​+ℓ=1∑L​wℓ​(Γ^ℓ​+Γ^ℓ⊤​),wℓ​=1−L+1ℓ​,(5)

with Bartlett weights wℓw_\ellwℓ​ and Γ^ℓ=T−1∑txte^te^t−ℓxt−ℓ⊤\hat\Gamma_\ell = T^{-1}\sum_t x_t\hat e_t\hat e_{t-\ell}x_{t-\ell}^{\top}Γ^ℓ​=T−1∑t​xt​e^t​e^t−ℓ​xt−ℓ⊤​ is consistent and positive semidefinite as L→∞L \to \inftyL→∞ with L/T→0L/T \to 0L/T→0, while the conventional σ2Q−1\sigma^2 Q^{-1}σ2Q−1 understates the variance when the ete_tet​ are positively autocorrelated. :::

::: proof Write T(β^−β)=(T−1∑xtxt⊤)−1 T−1/2∑xtet\sqrt{T}(\hat\beta - \beta) = (T^{-1}\sum x_t x_t^{\top})^{-1}\,T^{-1/2}\sum x_t e_tT​(β^​−β)=(T−1∑xt​xt⊤​)−1T−1/2∑xt​et​. The design average converges to QQQ by the law of large numbers, and the score average T−1/2∑xtetT^{-1/2}\sum x_t e_tT−1/2∑xt​et​ obeys a central limit theorem for the dependent sequence xtetx_t e_txt​et​, whose limiting variance is the long-run variance S=∑ℓ=−∞∞ΓℓS = \sum_{\ell=-\infty}^{\infty}\Gamma_\ellS=∑ℓ=−∞∞​Γℓ​ with Γ−ℓ=Γℓ⊤\Gamma_{-\ell} = \Gamma_\ell^{\top}Γ−ℓ​=Γℓ⊤​, giving the sandwich Q−1SQ−1Q^{-1} S Q^{-1}Q−1SQ−1. The Bartlett taper wℓ=1−ℓ/(L+1)w_\ell = 1 - \ell/(L+1)wℓ​=1−ℓ/(L+1) makes S^\hat SS^ a weighted sum of sample autocovariances that is positive semidefinite by construction and consistent under L/T→0L/T \to 0L/T→0 [6]. Setting all wℓ=0w_\ell = 0wℓ​=0 recovers the White heteroskedasticity-consistent estimator, valid when the ete_tet​ are uncorrelated but heteroskedastic [7], [8]. :::

Two routes then restore valid inference. Model ete_tet​ as an ARMA and estimate the mean and error parameters jointly by generalised least squares, the Cochrane-Orcutt approach, efficient under a correct error model, or keep ordinary least squares and report Equation (5), agnostic to it [9]. Residual serial correlation is checked by the Ljung-Box portmanteau statistic rather than the Durbin-Watson d=∑t(Δe^t)2/∑te^t2≈2(1−ρ^1)d = \sum_t(\Delta\hat e_t)^2/\sum_t\hat e_t^2 \approx 2(1 - \hat\rho_1)d=∑t​(Δe^t​)2/∑t​e^t2​≈2(1−ρ^​1​), which reads only the lag-one autocorrelation [10], [11].

The moving average, the ARMA model, and their long-memory and seasonal extensions are the linear vocabulary of serial dependence. The autocorrelation and partial autocorrelation identify the pure models by their dual truncation, the extended autocorrelation identifies the mixed one, and long division exposes the impulse responses that forecasting and the Wold representation both rest on. Where the geometric rate gives way to a unit root, the same polynomials open onto the unit-root and cointegration analysis, and where the conditional mean is adequate but the conditional variance is not, the squared residuals carry the autocorrelation that the level lacks [12], [13].

[1]
G. E. P. Box, G. M. Jenkins, and G. C. Reinsel, Time Series Analysis: Forecasting and Control, 3rd ed. Prentice Hall, 1994.
[2]
P. J. Brockwell and R. A. Davis, Time Series: Theory and Methods, 2nd ed. Springer, 1991.
[3]
R. S. Tsay and G. C. Tiao, “Consistent estimates of autoregressive parameters and extended sample autocorrelation function for stationary and nonstationary ARMA models,” Journal of the American Statistical Association, vol. 79, no. 385, pp. 84–96, 1984.
[4]
J. R. M. Hosking, “Fractional differencing,” Biometrika, vol. 68, no. 1, pp. 165–176, 1981.
[5]
Z. Ding, C. W. J. Granger, and R. F. Engle, “A long memory property of stock returns and a new model,” Journal of Empirical Finance, vol. 1, no. 1, pp. 83–106, 1993.
[6]
W. K. Newey and K. D. West, “A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix,” Econometrica, vol. 55, no. 3, pp. 703–708, 1987.
[7]
H. White, “A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity,” Econometrica, vol. 48, no. 4, pp. 817–838, 1980.
[8]
F. Eicker, “Limit theorems for regressions with unequal and dependent errors,” in Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, vol. 1, University of California Press, 1967, pp. 59–82.
[9]
W. H. Greene, Econometric Analysis, 5th ed. Prentice Hall, 2003.
[10]
G. M. Ljung and G. E. P. Box, “On a measure of lack of fit in time series models,” Biometrika, vol. 65, no. 2, pp. 297–303, 1978.
[11]
G. E. P. Box and D. A. Pierce, “Distribution of residual autocorrelations in autoregressive-integrated moving average time series models,” Journal of the American Statistical Association, vol. 65, no. 332, pp. 1509–1526, 1970.
[12]
W. A. Fuller, Introduction to Statistical Time Series. Wiley, 1976.
[13]
J. Y. Campbell, A. W. Lo, and A. C. MacKinlay, The Econometrics of Financial Markets. Princeton University Press, 1997.

Part 3 of 6 in Econometrics

← previousAutoregressive Modelsnext →Unit Roots and Stationarity Tests

Explore connections

see in the atlas →

related

  • Autocorrelation in Financial Time Series
  • Conditional Heteroskedasticity and GARCH
  • Cointegration
cite
@misc{serial-dependence,
  author = {Zac Kienzle},
  title  = {Serial Dependence: moving averages, ARMA, and long memory},
  year   = {2026},
  month  = {07},
  url    = {https://zackienzle.com/blog/serial-dependence}
}