Skip to content
homethesisprojectswritingaboutworkquestionsarithmeticsequencesresume
Loading
~/blog/cointegration0%dark
  1. home/
  2. writing/
  3. Cointegration

08 July 2026 · 4 min read · updated 23 Sept 2026

Cointegration

A single integrated series wanders without bound, yet a linear combination of several can be stationary, cancelling the common trends and leaving a mean-reverting spread. That is cointegration. We derive the error-correction representation as an exact reparametrisation of a vector autoregression, read off the cointegrating vectors and the speeds at which each equation corrects disequilibrium, prove the super-consistency of the Engle-Granger regression and the reduced-rank structure behind the Johansen test, and close with the threshold band of inaction that a transaction cost carves into the correction.

  • 3 equations
  • 7 connections
  • time-series
  • econometrics
  • stochastic-processes
On this page▾
  • Cointegration and error correction
  • Engle-Granger
  • Johansen
  • Threshold cointegration

4 min left

  • Cointegration and error correction2m
  • Engle-Granger1m
  • Johansen1m
  • Threshold cointegration1m

A single integrated series accumulates its shocks and drifts, but a linear combination of several integrated series can cancel the common trends and revert to a level. That combination is cointegration, and its stationary spread is the object a statistical arbitrage strategy trades. This post builds the error-correction representation from a vector autoregression, splits it into the cointegrating vectors and the adjustment speeds, and derives the two tests that certify the rank [1], [2].

#Cointegration and error correction

::: definition [def

] A series is I(0)I(0)I(0) when it is stationary with a positive finite long-run variance, and I(d)I(d)I(d) when its dddth difference is I(0)I(0)I(0) but its (d−1)(d-1)(d−1)th is not. An I(1)I(1)I(1) series has a unit root, accumulates shocks, and has variance growing linearly in ttt. :::

::: definition [def

] An I(1)I(1)I(1) vector yt∈Rny_t\in\mathbb{R}^nyt​∈Rn is cointegrated with rank rrr when there is an n×rn\times rn×r matrix β\betaβ of full rank with β⊤yt\beta^{\top}y_tβ⊤yt​ stationary. Each column of β\betaβ is a cointegrating vector, a linear combination that cancels the common trends and leaves a stationary spread. :::

The link to a vector autoregression is exact. Let yt=∑i=1pAiyt−i+εty_t=\sum_{i=1}^p A_i y_{t-i}+\varepsilon_tyt​=∑i=1p​Ai​yt−i​+εt​ with white noise εt\varepsilon_tεt​ and A(z)=In−∑iAiziA(z)=I_n-\sum_i A_i z^iA(z)=In​−∑i​Ai​zi.

::: theorem [thm

] The order-ppp autoregression reparametrises identically as

Δyt=Π yt−1+∑i=1p−1Γi Δyt−i+εt,Π=−A(1)=−(In−∑i=1pAi),Γi=−∑j=i+1pAj.(1)\Delta y_t=\Pi\,y_{t-1}+\sum_{i=1}^{p-1}\Gamma_i\,\Delta y_{t-i}+\varepsilon_t, \qquad \Pi=-A(1)=-\Bigl(I_n-\sum_{i=1}^{p}A_i\Bigr),\quad \Gamma_i=-\sum_{j=i+1}^{p}A_j. \tag*{(1)}Δyt​=Πyt−1​+i=1∑p−1​Γi​Δyt−i​+εt​,Π=−A(1)=−(In​−i=1∑p​Ai​),Γi​=−j=i+1∑p​Aj​.(1)

If yty_tyt​ is I(1)I(1)I(1) then rank⁡Π=r<n\operatorname{rank}\Pi=r<nrankΠ=r<n and Π=αβ⊤\Pi=\alpha\beta^{\top}Π=αβ⊤ with α,β\alpha,\betaα,β of full rank rrr; the columns of β\betaβ are the cointegrating vectors and the rows of α\alphaα the speeds at which each equation corrects past disequilibrium. When r=nr=nr=n the level is already stationary; when r=0r=0r=0 no stationary combination exists and the system is a pure unit-root vector. :::

::: proof Add and subtract to telescope the levels into differences. From yt=∑i=1pAiyt−i+εty_t=\sum_{i=1}^p A_i y_{t-i}+\varepsilon_tyt​=∑i=1p​Ai​yt−i​+εt​ subtract yt−1y_{t-1}yt−1​ and write each yt−i=yt−1−∑m=1i−1Δyt−my_{t-i}=y_{t-1}-\sum_{m=1}^{i-1}\Delta y_{t-m}yt−i​=yt−1​−∑m=1i−1​Δyt−m​. Collecting the coefficient of yt−1y_{t-1}yt−1​ gives ∑i=1pAi−In=−A(1)=Π\sum_{i=1}^p A_i-I_n=-A(1)=\Pi∑i=1p​Ai​−In​=−A(1)=Π, and collecting the coefficient of Δyt−i\Delta y_{t-i}Δyt−i​ gives −∑j=i+1pAj=Γi-\sum_{j=i+1}^{p}A_j=\Gamma_i−∑j=i+1p​Aj​=Γi​, which is Equation (1). Since Δyt\Delta y_tΔyt​, the Δyt−i\Delta y_{t-i}Δyt−i​, and εt\varepsilon_tεt​ are stationary, the term Πyt−1\Pi y_{t-1}Πyt−1​ must be stationary though yt−1y_{t-1}yt−1​ is I(1)I(1)I(1), so Π\PiΠ has reduced rank r=rank⁡Π<nr=\operatorname{rank}\Pi<nr=rankΠ<n. Take the rank factorisation Π=αβ⊤\Pi=\alpha\beta^{\top}Π=αβ⊤ with α,β\alpha,\betaα,β of full column rank rrr; left-multiplying Πyt−1=α(β⊤yt−1)\Pi y_{t-1}=\alpha(\beta^{\top}y_{t-1})Πyt−1​=α(β⊤yt−1​) by a left inverse of α\alphaα shows β⊤yt−1\beta^{\top}y_{t-1}β⊤yt−1​ is stationary, so the columns of β\betaβ are the rrr cointegrating vectors. :::

::: remark This is the necessary direction. The converse, that a rank-rrr Π\PiΠ makes yty_tyt​ exactly I(1)I(1)I(1) with rrr cointegrating relations and n−rn-rn−r common trends, is the Granger representation theorem and needs α⊥⊤Γβ⊥\alpha_{\perp}^{\top}\Gamma\beta_{\perp}α⊥⊤​Γβ⊥​ nonsingular, with Γ=In−∑i=1p−1Γi\Gamma=I_n-\sum_{i=1}^{p-1}\Gamma_iΓ=In​−∑i=1p−1​Γi​ and α⊥,β⊥\alpha_{\perp},\beta_{\perp}α⊥​,β⊥​ the orthogonal complements of α,β\alpha,\betaα,β, which rules out I(2)I(2)I(2) behaviour. :::

The factorisation separates levels from dynamics. The columns of β\betaβ are the stationary combinations, β⊤yt\beta^{\top}y_tβ⊤yt​ the disequilibrium each carries, and α\alphaα the rate at which each equation pulls back toward it. The directions orthogonal to α\alphaα carry what does not revert.

::: proposition [prop

] Let α⊥\alpha_\perpα⊥​ be a full-rank n×(n−r)n\times(n-r)n×(n−r) orthogonal complement of α\alphaα, so α⊥⊤α=0\alpha_\perp^{\top}\alpha=0α⊥⊤​α=0. Then α⊥⊤yt\alpha_\perp^{\top}y_tα⊥⊤​yt​ carries the n−rn-rn−r common stochastic trends of the system and is not mean reverting. :::

::: proof Premultiply Equation (1) by α⊥⊤\alpha_\perp^{\top}α⊥⊤​. The error-correction term drops because α⊥⊤α=0\alpha_\perp^{\top}\alpha=0α⊥⊤​α=0, leaving α⊥⊤Δyt=∑iα⊥⊤ΓiΔyt−i+α⊥⊤εt\alpha_\perp^{\top}\Delta y_t=\sum_i\alpha_\perp^{\top}\Gamma_i\Delta y_{t-i}+\alpha_\perp^{\top}\varepsilon_tα⊥⊤​Δyt​=∑i​α⊥⊤​Γi​Δyt−i​+α⊥⊤​εt​, a stationary series with no level feedback. Cumulating, α⊥⊤yt\alpha_\perp^{\top}y_tα⊥⊤​yt​ is a partial sum of those stationary increments, and under the Granger representation condition its long-run variance is nonsingular, so it is I(1)I(1)I(1) and supplies the n−rn-rn−r trends that the rrr stationary combinations β⊤yt\beta^{\top}y_tβ⊤yt​ do not span [3]. :::

#Engle-Granger

The simplest estimator fixes one spread by a regression in levels, and the unit-root theory of the previous post makes it converge unusually fast.

::: theorem [thm

] Let y0t=β⊤y1t+uty_{0t}=\beta^{\top}y_{1t}+u_ty0t​=β⊤y1t​+ut​ with y1ty_{1t}y1t​ a kkk-vector I(1)I(1)I(1) process and utu_tut​ stationary. The ordinary least squares estimator is super-consistent, T(β^−β)=Op(1)T(\hat\beta-\beta)=O_p(1)T(β^​−β)=Op​(1), converging at rate TTT rather than T1/2T^{1/2}T1/2. :::

::: proof β^−β=(∑y1ty1t⊤)−1∑y1tut\hat\beta-\beta=\bigl(\sum y_{1t}y_{1t}^{\top}\bigr)^{-1}\sum y_{1t}u_tβ^​−β=(∑y1t​y1t⊤​)−1∑y1t​ut​. The normalised design satisfies T−2∑y1ty1t⊤⇒∫01BB⊤T^{-2}\sum y_{1t}y_{1t}^{\top}\cvW\int_0^1 BB^{\top}T−2∑y1t​y1t⊤​⇒∫01​BB⊤ for a kkk-vector Brownian motion BBB, the matrix analogue of the first functional in the unit-root post, while T−1∑y1tut=Op(1)T^{-1}\sum y_{1t}u_t=O_p(1)T−1∑y1t​ut​=Op​(1). The ratio gives T(β^−β)⇒(∫BB⊤)−1 ⁣∫B dBuT(\hat\beta-\beta)\cvW\bigl(\int BB^{\top}\bigr)^{-1}\!\int B\,\dd B_uT(β^​−β)⇒(∫BB⊤)−1∫BdBu​, an Op(1)O_p(1)Op​(1) limit, so the error is of order T−1T^{-1}T−1. :::

::: remark Cointegration is tested by a Dickey-Fuller statistic on the residual u^t\hat u_tu^t​. Because β\betaβ is estimated, the null limit is not the Dickey-Fuller law but a residual functional whose quantiles depend on kkk, the Engle-Granger, equivalently Phillips-Ouliaris, critical values [4]. :::

Engle-Granger cointegration test
  input   I(1) series y0 and k-vector of I(1) regressors y1
  1  regress y0 on y1 by least squares, leaving residual u_hat
  2  apply the augmented Dickey-Fuller statistic to u_hat
     with no deterministic term
  3  reject no cointegration when the statistic is below the
     Phillips-Ouliaris critical value for k regressors

#Johansen

Engle-Granger fixes one spread by regression. Johansen estimates the rank rrr and the whole cointegrating matrix β\betaβ at once by Gaussian maximum likelihood on the error-correction system.

::: theorem [thm

] Concentrating out the lagged differences leaves residuals R0t,R1tR_{0t},R_{1t}R0t​,R1t​ with product moments Sij=T−1∑tRitRjt⊤S_{ij}=T^{-1}\sum_t R_{it}R_{jt}^{\top}Sij​=T−1∑t​Rit​Rjt⊤​. Gaussian maximum likelihood reduces to the ordered eigenvalues λ^1≥⋯≥λ^n\hat\lambda_1\ge\dots\ge\hat\lambda_nλ^1​≥⋯≥λ^n​ of ∣λS11−S10S00−1S01∣=0\lvert\lambda S_{11}-S_{10}S_{00}^{-1}S_{01}\rvert=0∣λS11​−S10​S00−1​S01​∣=0, and

LRtr(r)=−T ⁣ ⁣∑i=r+1n ⁣ln⁡(1−λ^i),LRmax⁡(r)=−Tln⁡(1−λ^r+1).(2)\mathrm{LR}_{\mathrm{tr}}(r)=-T\!\!\sum_{i=r+1}^{n}\!\ln(1-\hat\lambda_i),\qquad \mathrm{LR}_{\max}(r)=-T\ln(1-\hat\lambda_{r+1}). \tag*{(2)}LRtr​(r)=−Ti=r+1∑n​ln(1−λ^i​),LRmax​(r)=−Tln(1−λ^r+1​).(2)

Under the rank-rrr null LRtr(r)\mathrm{LR}_{\mathrm{tr}}(r)LRtr​(r) converges to a functional of an (n−r)(n-r)(n−r)-dimensional Brownian motion. :::

::: proof The λ^i\hat\lambda_iλ^i​ are squared sample canonical correlations between Δyt\Delta y_tΔyt​ and yt−1y_{t-1}yt−1​ after the lags are removed. The rrr largest are Op(1)O_p(1)Op​(1) and identify genuine cointegrating directions; the smallest n−rn-rn−r are Op(T−1)O_p(T^{-1})Op​(T−1), so −Tln⁡(1−λ^i)=Tλ^i+op(1)-T\ln(1-\hat\lambda_i)=T\hat\lambda_i+o_p(1)−Tln(1−λ^i​)=Tλ^i​+op​(1). On the non-cointegrated directions the Tλ^iT\hat\lambda_iTλ^i​ are eigenvalues of a matrix in the integrated regressors and the innovations, whose joint limit is the stated Brownian functional by the functional central limit theorem and the continuous-mapping theorem. :::

Johansen trace test for cointegrating rank
  input   series y, lag order, deterministic case
  1  concentrate out the lagged differences, leaving S00, S01, S11
  2  lambda_1 >= .. >= lambda_n <- eigenvalues of
     |lambda S11 - S10 S00^{-1} S01| = 0
  3  for r = 0, 1, .., n-1:
       LR_tr(r) <- -T sum_{i>r} ln(1 - lambda_i)
       if LR_tr(r) below its critical value, return rank r
  4  return full rank n

::: remark A constant or trend enters either inside the cointegrating space, where it sets the level or slope of β⊤yt\beta^{\top}y_tβ⊤yt​, or outside it, where it drives the common trends. Five nested cases are standard, from no deterministic term through a restricted constant, an unrestricted constant, a restricted trend, to an unrestricted trend, each with its own null limit for the rank statistics. Misreading a drift as an in-space mean inflates the apparent rank, so the test is run against the specification matching the data. :::

#Threshold cointegration

Linear error correction pulls back at a constant speed however far the system has strayed. When a band of inaction separates small departures from large ones, as a no-arbitrage region bounded by transaction costs does, the correction switches on only outside the band.

::: definition [def

] Let zt=β⊤ytz_t=\beta^{\top}y_tzt​=β⊤yt​ be the disequilibrium of a cointegrated system. The multivariate threshold error-correction model makes the dynamics regime-dependent in zt−1z_{t-1}zt−1​,

Δyt=c(j)+∑i=1pΦi(j)Δyt−i+α(j)zt−1+εt(j),j=j(zt−1),(3)\Delta y_t=c^{(j)}+\sum_{i=1}^{p}\Phi_i^{(j)}\Delta y_{t-i}+\alpha^{(j)}z_{t-1}+\varepsilon_t^{(j)},\qquad j=j(z_{t-1}), \tag*{(3)}Δyt​=c(j)+i=1∑p​Φi(j)​Δyt−i​+α(j)zt−1​+εt(j)​,j=j(zt−1​),(3)

the regime jjj chosen by which interval cut by thresholds γ1<γ2\gamma_1<\gamma_2γ1​<γ2​ contains zt−1z_{t-1}zt−1​. :::

::: remark With three regimes the outer two carry the correction while the middle one, γ1<zt−1≤γ2\gamma_1<z_{t-1}\le\gamma_2γ1​<zt−1​≤γ2​, has α(2)≈0\alpha^{(2)}\approx0α(2)≈0, so inside the band the spread drifts as an uncorrected random walk and outside it is pulled back. The band is the region where the gain from correcting does not cover its cost, and a test of α(2)=0\alpha^{(2)}=0α(2)=0 against a correcting outer regime is a test for threshold cointegration [5]. The thresholds are fixed by minimising the residual sum of squares over a grid. :::

Cointegration completes the unit-root picture. Where the unit-root tests certify that each series is individually I(1)I(1)I(1), the error-correction representation certifies that a combination of them is I(0)I(0)I(0), splitting the system into n−rn-rn−r common trends that wander and rrr stationary spreads that revert. That decomposition of a nonstationary vector into permanent and transitory parts is the econometric backbone of mean-reversion strategies such as statistical arbitrage, and it sits one step beyond the autoregression whose reduced-rank coefficient matrix it reads.

[1]
R. F. Engle and C. W. J. Granger, “Co-integration and error correction: representation, estimation, and testing,” Econometrica, vol. 55, no. 2, pp. 251–276, 1987.
[2]
S. Johansen, “Estimation and hypothesis testing of cointegration vectors in Gaussian vector autoregressive models,” Econometrica, vol. 59, no. 6, pp. 1551–1580, 1991.
[3]
J. H. Stock and M. W. Watson, “Testing for common trends,” Journal of the American Statistical Association, vol. 83, no. 404, pp. 1097–1107, 1988.
[4]
P. C. B. Phillips and S. Ouliaris, “Asymptotic properties of residual based tests for cointegration,” Econometrica, vol. 58, no. 1, pp. 165–193, 1990.
[5]
N. S. Balke and T. B. Fomby, “Threshold cointegration,” International Economic Review, vol. 38, no. 3, pp. 627–645, 1997.

Part 5 of 6 in Econometrics

← previousUnit Roots and Stationarity Testsnext →Conditional Heteroskedasticity and GARCH

Explore connections

see in the atlas →

related

  • Statistical Arbitrage
  • The Ornstein-Uhlenbeck Process
  • Autoregressive Models

referenced by (2)

  • Serial Dependence: moving averages, ARMA, and long memory
  • Unit Roots and Stationarity Tests
cite
@misc{cointegration,
  author = {Zac Kienzle},
  title  = {Cointegration},
  year   = {2026},
  month  = {07},
  url    = {https://zackienzle.com/blog/cointegration}
}