Cointegration: Proof Obligations
This page states, for each result the cointegration programme needs and that is not established, exactly what must be proven and why. It is the proof-target companion to Cointegration, which holds only established mathematics, and to the temporal-cointegration plan, which holds the programme and its implementation stages.
Nothing on this page is a result. Each item is an obligation: a precisely stated theorem to be proven, or refuted by an explicit counterexample, or shown to be ill-posed and replaced. An obligation is discharged in exactly one of three ways and is then migrated out of this page:
- a theorem with a complete proof, which moves to Cointegration as established mathematics;
- a counterexample with a proof, recorded as a negative result within the obligation, together with the corrected statement that replaces it;
- a well-posedness finding, showing the obligation as stated is not a well-defined mathematical question, with the reformulation.
An item is not discharged by a simulation consistent with it, by a proof sketch, or by an appeal to the stationary analogue. Where a line of the proof is routine and where the research content lies is stated for each item, so that effort is not spent on the routine part.
The setting and the symbols below are those of Cointegration and Stochastic Processes. All obligations are stated for \(N\) of order \(10^3\), aspect ratio \(\gamma = N/T\) of order one, and the observability ordering \(\Delta \ll h \ll D \le T\).
Setting and definitions
Fix a filtered probability space \((\Omega, \mathcal{F}, (\mathcal{F}_t)_{t \ge 0}, \mathbb{P})\). Time is continuous unless a discrete index is stated; observations are recorded at integer multiples of the sampling interval \(\Delta\).
Let \(\beta \in \mathbb{R}^{N \times r}\) be fixed of full column rank, and let \(\beta_\perp\) be a complement with \(\beta^\top \beta_\perp = 0\). The residual and trend coordinates are \(z_t = \beta^\top p_t\) and \(f_t = \beta_\perp^\top p_t\), and the reconstruction identity is \(p_t = \beta(\beta^\top\beta)^{-1} z_t + \beta_\perp(\beta_\perp^\top\beta_\perp)^{-1} f_t\).
The regime process \(s_t \in \lbrace 0, 1 \rbrace^r\) is a semi-Markov process: it changes value at stopping times \(\tau_0 < \tau_1 < \dots\), the sojourn in state \(u\) has law \(F_u\) supported on \([D_{\min}, \infty)\) with finite mean \(m_u\) and \(F_u(D_{\min}) = 0\), and the embedded chain has transition matrix \(R\) with \(R_{uu} = 0\). Write \(\pi = \lim_{t\to\infty} \mathbb{E}[s_t]\) for the stationary activation vector, assumed to exist.
The residual follows the switching Ornstein-Uhlenbeck dynamics
\[ d z_t = \operatorname{diag}(s_t)\, \Theta (\mu - z_t)\, dt + \Sigma_z\, dW_t, \qquad \Theta = \theta I_r, \quad \Sigma_z \Sigma_z^\top \succ 0, \]
so that when \(s_{j,t} = 1\) the relation \(j\) mean-reverts with half-life \(h_j = \ln 2 / \theta_j\), and when \(s_{j,t} = 0\) it is a driftless random walk. The trend is a driftless random walk, \(f_t = f_0 + \Sigma_f B_t\), with \(W, B\) independent standard Brownian motions.
The observation layer maps the latent state to the data, \(y_t = \mathcal{Q}\big(p_t + \eta_t\big)\), where \(\eta_t\) is microstructure noise and \(\mathcal{Q}\) is rounding to a tick grid. During the theoretical programme the layer is switched off until an obligation concerns it.
Two definitions are needed and neither is standard. Call \(p\) interval-cointegrated with activity \(s\) and vectors \(\beta\) if, on every maximal interval on which \(s\) is constant, the restriction of \(\beta^\top p\) to that interval is stationary. The local rank at time \(t\) is \(r_t = \sum_{j=1}^{r} s_{j,t}\), the number of active relations; the ergodic rank is the rank of the ergodically averaged error-correction matrix \(\bar\Pi = \lim_{T\to\infty} T^{-1}\sum_{t} \alpha \operatorname{diag}(s_t)\beta^\top\).
P1. Well-posedness and the switching representation
Statement. Prove the following, which jointly make interval-cointegration a defensible process class.
- For every regime path and every initial condition, the switching residual SDE has a pathwise unique strong solution, and the reconstruction \(p_t = \beta(\beta^\top\beta)^{-1} z_t + \beta_\perp(\beta_\perp^\top\beta_\perp)^{-1} f_t\) is continuous at every switching instant.
- On each maximal active interval, the law of \(z\) converges to the Ornstein-Uhlenbeck stationary law as the interval length grows, uniformly in the initial condition, at the rate governed by \(\exp(-\theta D_{\min})\).
- The open direction. Every \(I(1)\) process that is interval-cointegrated with activity \(s\) admits a representation of the form above, with \(\alpha_t = \alpha \operatorname{diag}(s_t)\), a short-run part, and \(\alpha_\perp^\top \Gamma \beta_\perp\) invertible on each active interval.
Why it is required. Without part 3, "interval-cointegrated" is a definition with no representation theorem behind it, and no estimator has a stated target. Part 3 is the converse of the construction, in the same position that the Granger representation theorem occupies for the contiguous case.
What would refute it. A process that is interval-cointegrated yet admits no error-correction representation on some active interval, or a process with well-defined local behaviour but no well-defined path at a switching instant.
Strategy and division of labour. Parts 1 and 2 are routine: Lipschitz drift and the explicit integrating factor give both, and part 2 is the standard Ornstein-Uhlenbeck convergence with the switching entering only through the interval length. Part 3 is the research content; it requires a Beveridge-Nelson-type decomposition adapted to a random partition of the time axis.
Depends on. Nothing.
P2. Identifiability and consistency of rank under switching
Statement. Let the regime chain be ergodic and irreducible on its state space, so \(\pi_j > 0\) for every \(j\), and let the Johansen rank estimator \(\hat r_T\) be the number of eigenvalues of \(S_{11}^{-1} S_{10} S_{00}^{-1} S_{01}\) above a threshold \(c_T \to 0\). Prove:
- \(\operatorname{rank}(\bar\Pi) = r\); equivalently, the \(j\)-th adjustment column is attenuated by the factor \(\pi_j\), so rank is identified if and only if \(\pi_j > 0\) for all \(j\) and \(\alpha_j \neq 0\).
- \(\hat r_T \to r\) in probability.
- The open direction. The limiting law of \(T(\hat r_T - r)\) is a mixture of the contiguous Johansen functional, with mixing weights determined by the stationary distribution of the regime chain and the sojourn laws \(F_u\). Give the mixture explicitly.
- The boundary. If some \(F_u\) is heavy-tailed so that active intervals have vanishing density, consistency fails. Give the precise tail condition that separates consistency from failure.
Why it is required. It decides whether the tabulated contiguous critical values may be used at all, and it is the first place where switching changes the answer rather than merely the sampling variability. Part 1 also gives the correct interpretation of an estimated loading: it is the true loading scaled by the activation probability, so a small \(\hat\alpha_j\) does not distinguish a weak relation from a rarely active one.
What would refute it. A regime law with \(\pi_j > 0\) for which \(\hat r_T\) is inconsistent, which would show that positive activation probability is not sufficient for identifiability.
Strategy and division of labour. Part 1 is routine: an ergodic theorem gives \(T^{-1}\sum_t \operatorname{diag}(s_t) \to \operatorname{diag}(\pi)\) and rank is preserved by invertible diagonal scaling. Parts 2 and 4 are the research content, requiring a functional central limit theorem for a triangular array whose regressor distribution is randomly time-varying. Part 3 requires characterizing the limit of an averaged quadratic form in matrix Brownian motion under a random time change.
Depends on. P1.
P3. Exact null law of the joint time-and-subspace detector
Statement. Define the joint detector
\[ \Lambda_T = \sup_{\tau \in [\epsilon T, (1-\epsilon) T]} \; \sup_{\beta \in G_{r,N}} \; \mathrm{LR}\big(\tau, \beta\big), \]
where \(G_{r,N}\) is the Grassmannian of \(r\)-dimensional subspaces of \(\mathbb{R}^N\) and \(\mathrm{LR}(\tau, \beta)\) is the likelihood-ratio statistic for a rank-\(r\) relation with vectors \(\beta\) active on a neighbourhood of \(\tau\). Under the null that \(p\) is \(I(1)\) with rank \(0\), derive the limiting law of \(\Lambda_T\) after the appropriate centering and scaling.
The conjectured limit object is a functional of a two-parameter process indexed by (time, subspace), the subspace parameter living on \(G_{r,N}\), obtained as the supremum of a quadratic form in matrix Brownian motion over the product of the time interval and the Grassmannian. Establish whether the two suprema commute in the limit.
Why it is required. This is the exact solution of the detection half. A detector that searches over both change times and subsets has a null distribution that is neither the Johansen law nor the single-break law; without it, the threshold is a simulated table for one parameter setting and cannot be transported.
What would refute it. A proof that \(\Lambda_T\) has no limiting law under any consistent scaling, together with the scaling that restores one, or a proof that the limit is degenerate.
Strategy and division of labour. The one-dimensional suprema each have known limits; the obligation is the interaction. The starting point is a functional central limit theorem for the score process indexed by \(\tau\) uniformly over \(G_{r,N}\), followed by a continuous-mapping argument on the product space, with care that \(G_{r,N}\) is non-compact in \(N\).
Depends on. P2, P7.
P4. Random matrix theory for the rank problem with \(I(1)\) levels
Statement. Let \(S_{ij} = T^{-1} \sum_t R_{it} R_{jt}^\top\) be the matrices of the reduced-rank regression, built from the \(I(1)\) levels. Prove, or disprove, a functional limit theorem for the spectrum of
\[ M_T = S_{11}^{-1/2} S_{10} S_{00}^{-1} S_{01} S_{11}^{-1/2} \]
as \(N, T \to \infty\) with \(N/T \to \gamma \in (0, \infty)\). Determine whether the largest eigenvalue of \(M_T\) under the null of rank zero separates from a bulk, and if so give the bulk law, the edge, and the fluctuation class. Define the cointegration threshold \(\ell_c^{\mathrm{coint}}\) as the sharp spike strength at which a rank-one alternative becomes detectable, and establish whether \(\ell_c^{\mathrm{coint}} = \sqrt{\gamma}\), or \(\ell_c^{\mathrm{coint}} > \sqrt{\gamma}\), or \(\ell_c^{\mathrm{coint}} < \sqrt{\gamma}\).
Why it is required. Marchenko-Pastur and Baik-Ben Arous-Peche require independent or weakly mixing entries with a growing sample. The matrices here are built from the levels, which are \(I(1)\) and non-mixing, so those results do not apply as stated. Until this is resolved, no threshold for the 1000-dimensional rank problem is justified, and the comparison with the stationary threshold is the quantitative measure of how much harder the non-stationary problem is.
What would refute it. An argument that no eigenvalue separates from any bulk in this scaling, which would mean that rank is not recoverable at any signal strength when \(\gamma\) is of order one. That is itself a decisive negative result.
Strategy and division of labour. The obstruction is that the residual matrices involve \(T^{-1}\sum_t p_{t-1} \varepsilon_t^\top\), a functional of matrix Brownian motion, so the standard Wishart representation fails. The candidate route is to express \(M_T\) as a functional of an \(N \times T\) array with a non-stationary, non-mixing column structure and to seek a limit theorem for such functionals; this is the least developed of the obligations.
Depends on. P7.
P5. Sparse-spike thresholds and delocalization
Statement. Let \(\beta\) have support of size \(s \ll N\), so that the alternative is a sparse spike. Prove an upper and a lower bound for the minimal spike strength at which the support of \(\beta\) is recoverable, with the aspect ratio \(\gamma\) fixed. Determine whether the threshold scales as \(\sqrt{s \log N / T}\) or as \(\sqrt{N/T}\), and characterize the region of \((s, \gamma)\) in which the estimated eigenvector is delocalized and therefore does not identify the support.
Why it is required. Sparse \(\beta\) is the only design that answers the question of which dimensions are cointegrated, and the delocalized threshold is the wrong one to use for it. This obligation also fixes what "which dimensions" can mean: if the eigenvector is delocalized, the support is not recoverable even when the relation is detected.
What would refute it. A proof that sparse structure gives no improvement over the dense threshold in this non-stationary setting.
Depends on. P4.
P6. The detection floor as a minimax lower bound
Statement. Let \(\mathcal{P}(\ell)\) be the class of laws under which the system carries a rank-one relation of spike strength \(\ell\), and \(\mathcal{P}_0\) the rank-zero null. Prove that there is a constant \(c > 0\) such that for \(\ell < \ell_c\),
\[ \inf_{\hat r} \; \sup_{P \in \mathcal{P}(\ell) \cup \mathcal{P}_0} \; P\big(\hat r \text{ incorrect}\big) \;\ge\; c, \]
the infimum over all tests, and construct an estimator whose risk vanishes for \(\ell > \ell_c\). The \(\ell_c\) is that of P4, or of P5 in the sparse case.
Why it is required. It converts "our detector failed on this instance" into "no detector can succeed on this class", which is the only statement that closes the question of how long and how strongly a relation must be active. It is the mathematical content behind the claim that there is a hard floor.
What would refute it. An estimator that beats the conjectured floor, which would invalidate P4 or P5.
Strategy and division of labour. A Le Cam or Fano argument over a packing of the parameter space by alternative cointegrating vectors, with the divergence between the resulting laws controlled by the spike strength and the effective sample size per active interval, which is where the switching enters.
Depends on. P4, P7.
P7. Sharp timescale conditions
Statement. Turn the ordering \(\Delta \ll h \ll D\) into necessary and sufficient conditions. Let \(\alpha \in (0,1)\) be the test size and \(1-\beta\) the prescribed power. Prove:
- Resolvability. Distinguishing \(\theta = 0\) from \(\theta > 0\) at a fixed span requires \(\theta \Delta\) bounded away from zero, and the Fisher information of the discretely sampled experiment is \(\Theta\big((\theta \Delta)^2\big)\). Give the exact constant.
- Observability. For a fixed power, \(D / h \ge c(\alpha, \beta)\); give \(c(\alpha, \beta)\) and show it is necessary, not merely sufficient.
- Estimability. For a fixed power with detection and confirmation on disjoint sub-intervals, \(D / W \ge c'(\alpha, \beta)\), where \(W\) is the detection window.
Why it is required. These constants are what make the timescale ordering operational. Without them, every parameter choice in a generator or a study is undefended, and a failure cannot be attributed to the theory rather than to the implementation.
What would refute it. A demonstration that the conditions are sufficient but not necessary, with the correct necessary condition replacing each.
Strategy and division of labour. Local asymptotic normality for the active-interval experiment gives the Fisher information of part 1; contiguity across a switching boundary gives the loss from part 3. Part 2's constant is a power calculation under the exact Gaussian law of Cointegration and is the most tractable of the three.
Depends on. Nothing.
P8. Exact solvability class of the switching stopping problem
Statement. Consider the optimal stopping problem for the spread with regime switching, whose value on the continuation region solves the coupled system \(\mathcal{L}_k u_k = 0\) across regimes \(k\), with regime-dependent speed \(\theta_k\), volatility \(\sigma_k\), cost \(c_k\), and transition rates \(q_{kl}\). Characterize the set of \((\theta_k, \sigma_k, c_k, q_{kl})\) for which the coupled system admits a closed-form solution in the parabolic cylinder or confluent hypergeometric family, and prove that outside this set no such closure exists.
Why it is required. It defines which detected relations can be traded with an exact policy and where the boundary of the exact-solvable envelope lies. It is the same structural question asked of market making in multi-asset-market-making.md, so a shared answer is a structural result rather than a model-specific one.
What would refute it. A proof that no nontrivial switching structure closes, which fixes the envelope boundary as a negative result.
Depends on. P7.
P9. Exact likelihood under tick quantization and asynchronicity
Statement. With the observation layer active, the exact likelihood is \(\prod_t \Pr\big(y_t \in \text{bin} \mid z_{t-1}\big)\), a product of transition probabilities of the latent process between quantization bins. Prove whether these probabilities admit an exact (spectral, Hermite-function) representation in the multivariate switching case, as they do for a univariate Ornstein-Uhlenbeck with two barriers, or prove that they do not and give the correct approximation class and its error.
Why it is required. Every threshold derived under continuous synchronous observation is an upper bound on real performance, and this obligation quantifies the gap. It is also the only obligation in which the HFT observation layer, and not the latent process, is the object of study.
What would refute it. A proof of intractability, with the approximating class that replaces the exact likelihood.
Depends on. P1.
P10. Bias of the half-life estimator under switching
Statement. Let \(\hat\theta\) be the estimator obtained by applying the AR(1) estimator of Cointegration to the active sub-intervals only. Derive \(\mathbb{E}[\hat\theta] - \theta\) to leading order in \(T\), as a function of the activation vector \(\pi\), the sojourn laws \(F_u\), and the sampling ratio \(\theta \Delta\), and construct a bias-corrected estimator with vanishing bias.
Why it is required. The half-life determines tradeability. The contiguous bias is of order \(1/T\); if the switching bias is of larger order, half-life estimation is unreliable at the horizons the programme targets, and the tradeability filter would be systematically wrong.
What would refute it. A demonstration that the bias is not \(O(1/T)\) but of larger order, which would make the estimator unreliable and require a different estimator rather than a correction.
Depends on. P2.
Dependencies and discharge order
graph TD
P7[P7 Timescale conditions] --> P3[P3 Joint null law]
P7 --> P4[P4 RMT for I(1) levels]
P7 --> P6[P6 Minimax floor]
P1[P1 Representation] --> P2[P2 Rank identifiability]
P1 --> P9[P9 Quantized likelihood]
P2 --> P3
P2 --> P10[P10 Half-life bias]
P4 --> P5[P5 Sparse thresholds]
P4 --> P6
P7 --> P8[P8 Switching stopping]
The rule is that an obligation is started only after the obligations it depends on are discharged. The obligations that can start immediately are P1 and P7, which is why they are the first two in any schedule. P2 and P4 are the load bearing results: P2 decides whether rank is identifiable at all under switching, and P4 decides whether it is recoverable in high dimension.
What a discharged obligation looks like
For illustration of the required standard, a discharged obligation would read as a statement of the form: under assumptions A1 to Ak, for all parameter values in a specified set, a specified estimator satisfies a specified limit, with the limit object named and its distribution or value given, together with the reduction to the corresponding contiguous result as the sojourn laws degenerate to a point mass at \(D_{\min} \to \infty\). The reduction is the check that the statement is not vacuously general, and it is the analogue of the specializations listed in the validation targets of Cointegration.