Exact Solutions in Stochastic Optimal Control

This page derives the exact solutions to the stochastic optimal control problems used as validation targets: the Avellaneda-Stoikov market-making family (base, risk-neutral, drift, impact, and stationary), the Merton portfolio, and the linear-quadratic regulator. Each problem is stated with its dynamics, objective, and exact value function and policy. The problems without a closed form (the Heston and Hawkes market-making extensions and the American put) are on SOC Models Without Exact Solutions. The pure (control-free) PDE benchmarks — parabolic and elliptic — are on the PDE page; the underlying processes are on the processes page.

The general framework follows [@pham2009continuous, @fleming2006controlled, @yong1999stochastic].

Market making (Avellaneda-Stoikov)

Notation

SymbolMeaning
\(\gamma\)CARA risk-aversion coefficient
\(\sigma\)Mid-price volatility (constant in the base model)
\(k\)Fill-intensity decay, \(\lambda(\delta) = A e^{-k\delta}\)
\(A\)Base order-arrival intensity
\(\phi\)Running inventory penalty, \(\phi\, q^2\) per unit time
\(q \in \mathbb{Z}\)Inventory, truncated to \([-q_{\max}, q_{\max}]\)
\(\theta(t, q)\)Reduced CARA value, \(V = -e^{-\gamma(X + qS + \theta)}\)
\(v_q(t)\)Log-transform \(v_q = e^{k\,\theta(t,q)}\)
\(\tau\)Time to maturity \(\tau = T - t\)

Setting

The market maker controls the bid and ask half-spreads \(\delta_t^b, \delta_t^a\) relative to the mid-price \(S_t\):

\[S_t^{\text{bid}} = S_t - \delta_t^b, \qquad S_t^{\text{ask}} = S_t + \delta_t^a.\]

The state is \([S, q, X]\), with mid-price \(S\), inventory \(q\), and cash \(X\), driven by

\[ \begin{aligned} dS_t &= \mu(t, S_t)\,dt + \sigma(t, S_t)\,dW_t, \\ dq_t &= dN_t^b - dN_t^a, \\ dX_t &= (S_t + \delta_t^a)\,dN_t^a - (S_t - \delta_t^b)\,dN_t^b, \end{aligned} \]

where \(N_t^b, N_t^a\) are point processes with intensities \(\lambda^b(\delta^b)\) and \(\lambda^a(\delta^a)\). The standard specification is exponential,

\[\lambda^b(\delta^b) = A e^{-k\delta^b}, \qquad \lambda^a(\delta^a) = A e^{-k\delta^a}.\]

The objective is the Bolza-form expected terminal utility with a running inventory penalty \(\phi\, q^2\):

\[V(t, S, q, X) = \sup_{\delta^a, \delta^b} \mathbb{E}\!\left[ U\!\left(X_T + q_T S_T - \int_t^T \phi\, q_s^2\, ds\right) \;\middle|\; S_t=S, q_t=q, X_t=X\right].\]

CARA separation and reduced PDE

With CARA utility \(U(x) = -e^{-\gamma x}\) the value separates through the ansatz

\[V(t, S, q, X) = -\exp\!\big(-\gamma (X + qS + \theta(t, q))\big),\]

reducing the problem to a PDE in \(\theta(t,q)\) alone. This reduction is a translation symmetry of the state: the dynamics and terminal payoff depend on the mid-price \(S\) and cash \(X\) only through the mark-to-market wealth \(Y = X + qS\), so the value is invariant under the one-parameter shift \((S, X) \mapsto (S + c,\; X - qc)\) for any constant \(c\). The value therefore depends on \(S\) and \(X\) through \(Y\) alone, and the CARA terminal condition forces the exponential factor in \(Y\), leaving only the inventory correction \(\theta(t,q)\). To see this, set \(\mu = 0\), constant \(\sigma\), and exponential intensities. The HJB is [@bellman1952theory, @bellman1966dynamic]

\[ 0 = \partial_t V + \tfrac12 \sigma^2 \partial_S^2 V + \sup_{\delta^a} A e^{-k\delta^a}\big[V(t,S,q-1,X+S+\delta^a) - V\big] + \sup_{\delta^b} A e^{-k\delta^b}\big[V(t,S,q+1,X-S+\delta^b) - V\big]. \]

Substituting the ansatz, the price-diffusion term yields the quadratic inventory penalty

\[ \partial_t V = -\gamma\,(\partial_t\theta)\,V, \qquad \tfrac12\sigma^2 \partial_S^2 V = -\tfrac12\gamma^2\sigma^2 q^2\,V, \]

and each jump term factors across the state shift. For the ask side,

\[ V(t,S,q-1,X+S+\delta^a) = -e^{-\gamma(X + qS + \delta^a + \theta(t,q-1))} = V\, e^{-\gamma(\delta^a + \theta_{q-1} - \theta_q)}. \]

Writing \(\Delta_a = \theta(t,q-1) - \theta(t,q)\) and \(\Delta_b = \theta(t,q+1) - \theta(t,q)\), the \(-\gamma V\) factor cancels and the HJB reduces to an equation in \(\theta\) alone:

\[ \partial_t\theta - \tfrac12\gamma\sigma^2 q^2 + \sup_{\delta^a} \frac{A e^{-k\delta^a}}{\gamma}\big[1 - e^{-\gamma(\delta^a + \Delta_a)}\big] + \sup_{\delta^b} \frac{A e^{-k\delta^b}}{\gamma}\big[1 - e^{-\gamma(\delta^b + \Delta_b)}\big] = 0. \]

Each supremum is a scalar optimization. Differentiating \(f(\delta) = e^{-k\delta}\big[1 - e^{-\gamma(\delta + \Delta)}\big]\) and setting the derivative to zero gives \(k = (k+\gamma)\, e^{-\gamma(\delta^* + \Delta)}\), hence the optimal spreads

\[ \delta^{a*} = \frac{1}{\gamma}\ln\!\Big(1 + \frac{\gamma}{k}\Big) - \Delta_a = \frac{1}{\gamma}\ln\!\Big(1 + \frac{\gamma}{k}\Big) + \theta(t,q) - \theta(t,q-1), \]

\[ \delta^{b*} = \frac{1}{\gamma}\ln\!\Big(1 + \frac{\gamma}{k}\Big) - \Delta_b = \frac{1}{\gamma}\ln\!\Big(1 + \frac{\gamma}{k}\Big) + \theta(t,q) - \theta(t,q+1). \]

The optimal intensities are \(\lambda^{a*} = A e^{-k\delta^{a*}}\) and \(\lambda^{b*} = A e^{-k\delta^{b*}}\). At optimality \(e^{-\gamma(\delta^* + \Delta)} = \frac{k}{k+\gamma}\), so \(1 - e^{-\gamma(\delta^* + \Delta)} = \frac{\gamma}{k+\gamma}\) and the optimized Hamiltonian collapses to

\[H^* = \frac{\lambda^{a*} + \lambda^{b*}}{\gamma + k}.\]

Substituting back yields the reduced PDE for \(\theta\) (with the running penalty \(\phi\) restored):

\[ \partial_t\theta - \big(\tfrac12\gamma\sigma^2 + \phi\big) q^2 + \frac{A}{k+\gamma}\Big(1 + \frac{\gamma}{k}\Big)^{-k/\gamma} \Big[e^{-k(\theta_q - \theta_{q-1})} + e^{-k(\theta_q - \theta_{q+1})}\Big] = 0, \]

with terminal condition \(\theta(T, q) = 0\) (Mayer form). For the base model \(\phi = 0\).

Finite-horizon exact solution (matrix exponential)

Define \(v_q(t) = \exp(k\,\theta(t,q))\). The nonlinear PDE becomes the linear ODE system

\[ \dot{v}_q(t) = \tilde{\alpha}\, q^2 v_q(t) - \eta\big(v_{q-1}(t) + v_{q+1}(t)\big), \]

with constants

\[ \tilde{\alpha} = \tfrac{k}{2}\gamma\sigma^2 + k\phi, \qquad \eta = \frac{kA}{k+\gamma}\Big(1 + \frac{\gamma}{k}\Big)^{-k/\gamma}. \]

For the base model \(\phi = 0\), so \(\tilde{\alpha} = \tfrac{k}{2}\gamma\sigma^2\), and the terminal condition is \(v_q(T) = 1\).

Stacking \(v_q\) over \(q \in [-q_{\max}, q_{\max}]\) gives a linear system \(\dot{\mathbf{v}}(t) = G\,\mathbf{v}(t)\) with tridiagonal generator \(G_{ii} = \tilde{\alpha} q_i^2\), \(G_{i,i\pm1} = -\eta\). The solution is expressed in terms of the negated generator \(M = -G\), with

\[ M_{i,j} = \begin{cases} -\tilde{\alpha}\, q_i^2, & i = j, \\ +\eta, & |i-j| = 1, \\ 0, & \text{otherwise}, \end{cases} \]

as

\[ \mathbf{v}(t) = e^{M\,(T-t)}\,\mathbf{v}(T). \]

This negated form is the sign convention used throughout: the operator has negative diagonal and positive off-diagonal, and the matrix exponential is evaluated directly in \(T - t\).

Optimal spreads

From \(v_q(t)\), the optimal half-spreads are

\[ \begin{aligned} \delta^{a*}(t, q) &= \delta_0 + \frac{1}{k}\ln\!\Big(\frac{v_q(t)}{v_{q-1}(t)}\Big), \\ \delta^{b*}(t, q) &= \delta_0 + \frac{1}{k}\ln\!\Big(\frac{v_q(t)}{v_{q+1}(t)}\Big), \end{aligned} \qquad \delta_0 = \frac{1}{\gamma}\ln\!\Big(1 + \frac{\gamma}{k}\Big), \]

and the fill intensities are \(\lambda^{a*} = A e^{-k\delta^{a*}}\), \(\lambda^{b*} = A e^{-k\delta^{b*}}\). The base spread \(\delta_0\) is the asymptotic value at \(q = 0\) as \(t \to -\infty\); inventory skew enters through the \(\ln(v_q/v_{q\pm1})\) terms. As \(\gamma \to 0\) the base spread tends to \(1/k\), recovering the risk-neutral (linear-utility) solution below.

Value recovery

\[\theta(t,q) = \frac{1}{k}\ln v_q(t), \qquad V(t,S,q,X) = -\exp\!\big(-\gamma(X + qS + \theta(t,q))\big).\]

Terminal conditions

Terminal condition\(v_q(T)\)\(\theta(T,q)\)
Zero\(1\)\(0\)
Liquidation cost\(\exp\big(-\gamma\,q

Under the liquidation-cost condition, the terminal value penalizes nonzero inventory by the base spread, matching the cost of crossing the spread to unwind.

Risk-neutral (linear utility) reduction

Under linear (risk-neutral) utility \(U(x) = x\) the objective is the expected terminal mark-to-market wealth,

\[V(t, S, q, X) = \sup_{\delta^a, \delta^b} \mathbb{E}\big[X_T + q_T S_T \;\big|\; S_t=S, q_t=q, X_t=X\big].\]

The value separates as \(V = X + qS + \theta(t, q)\), the same translation symmetry with an affine (translation-invariant) rather than exponential value. Because \(V\) is affine in \(S\), the price-diffusion term \(\tfrac12\sigma^2\,\partial_S^2 V\) vanishes and with it the inventory-risk penalty. The HJB reduces to

\[0 = \partial_t\theta + \sup_{\delta^a} A e^{-k\delta^a}\big(\delta^a + \Delta_a\big) + \sup_{\delta^b} A e^{-k\delta^b}\big(\delta^b + \Delta_b\big),\]

with \(\Delta_a = \theta(t,q-1) - \theta(t,q)\) and \(\Delta_b = \theta(t,q+1) - \theta(t,q)\) as before. The first-order condition of \(f(\delta) = e^{-k\delta}(\delta + \Delta)\) is \(k(\delta^* + \Delta) = 1\), giving the optimal spreads

\[\delta^{a*} = \frac1k + \theta(t,q) - \theta(t,q-1), \qquad \delta^{b*} = \frac1k + \theta(t,q) - \theta(t,q+1),\]

with base spread \(\delta_0 = 1/k\). The log-linearization \(v_q = e^{k\,\theta(t,q)}\) turns the reduced equation into the linear system

\[\dot{v}_q(t) = -A e^{-1}\big(v_{q-1}(t) + v_{q+1}(t)\big),\]

with no diagonal term, or in the negated form

\[M_{i,j} = \begin{cases} +A e^{-1}, & |i-j| = 1, \\ 0, & \text{otherwise}, \end{cases} \qquad \mathbf{v}(t) = e^{M(T-t)}\,\mathbf{v}(T),\]

with \(v_q(T) = 1\) and value recovery \(V = X + qS + \tfrac1k\ln v_q\). This is the \(\gamma \to 0\) limit of the CARA solution: the diagonal \(\tilde\alpha = \tfrac{k}{2}\gamma\sigma^2\) vanishes and the off-diagonal \(\eta = \tfrac{kA}{k+\gamma}(1+\gamma/k)^{-k/\gamma}\) tends to \(A e^{-1}\), while the base spread \(\tfrac1\gamma\ln(1+\gamma/k)\) tends to \(1/k\).

Drift extension

A constant mid-price drift \(dS_t = \mu\,dt + \sigma\,dW_t\) adds a linear \(-k\mu q\) term to the diagonal of the generator, so in the negated form

\[ M_{i,j} = \begin{cases} -\tilde{\alpha}\, q_i^2 + k\mu\, q_i, & i = j, \\ +\eta, & |i-j| = 1. \end{cases} \]

The optimal target inventory is \(q^* = \dfrac{k\mu}{2\tilde{\alpha}}\); the drift breaks the \(q \mapsto -q\) reflection symmetry of the base generator, shifting the target away from zero.

Permanent-impact extension

A permanent price impact \(\xi\) shifts the mid-price by \(\pm\xi\) on each fill, \(dS_t = \sigma\,dW_t + \xi\,dN_t^a - \xi\,dN_t^b\), producing asymmetric off-diagonals and a nontrivial terminal condition:

\[ M_{i,j} = \begin{cases} -\tilde{\alpha}\, i^2, & j = i, \\ +\eta\, e^{k\xi(i-1)}, & j = i-1, \\ +\eta\, e^{-k\xi(i+1)}, & j = i+1, \end{cases} \qquad v_q(T) = \exp\!\Big(-\tfrac{k}{2}\xi\, q^2\Big). \]

The impact \(\xi\) breaks the reflection symmetry \(q \mapsto -q\) of the base generator: the up- and down-couplings \(e^{k\xi(i-1)}\) and \(e^{-k\xi(i+1)}\) differ once \(\xi \neq 0\), reflecting that a fill moves the mid-price in the direction of the trade. The optimal spreads gain impact adjustments

\[ \begin{aligned} \delta^{a*} &= \delta^{a*}_{\text{base}} - \xi(q - 1), \\ \delta^{b*} &= \delta^{b*}_{\text{base}} + \xi(q + 1). \end{aligned} \]

Stationary (infinite-horizon) solution

As \(T \to \infty\) the finite-horizon solution concentrates on the principal (Perron) eigenvector of \(M\). The undiscounted HJB is translation invariant and has no unique solution of \(\sup_u\lbrace f + \mathcal{L}^u V\rbrace = 0\); its exact limit is the eigenpair

\[M\,\mathbf{v} = \lambda_{\max}\,\mathbf{v},\]

where \(M\) is the symmetric tridiagonal operator with diagonal \(-\tilde{\alpha}\, q^2\) and off-diagonal \(+\eta\) defined above.

Gauge fixing

The eigenvector is defined up to a positive scalar, a scale symmetry \(\mathbf{v} \mapsto c\,\mathbf{v}\) of the eigenproblem \(M\mathbf{v} = \lambda_{\max}\mathbf{v}\). The gauge is fixed by normalizing

\[\theta(0) = 0 \quad\Longleftrightarrow\quad \theta(q) = \frac{1}{k}\ln\frac{v_q}{v_0}.\]

Stationary spreads

\[ \begin{aligned} \delta^{b*}(q) &= \delta_0 + \theta(q) - \theta(q+1), \\ \delta^{a*}(q) &= \delta_0 + \theta(q) - \theta(q-1). \end{aligned} \]

Perron-Frobenius guarantees

The operator \(M\) has non-negative off-diagonals and is irreducible on the truncated inventory grid, so by Perron-Frobenius:

  • the principal eigenvalue \(\lambda_{\max}\) is real and simple;
  • the principal eigenvector \(\mathbf{v}\) is strictly positive;
  • \(\theta(q) = \tfrac1k\ln(v_q/v_0)\) is even in \(q\) and decreasing in \(|q|\), since the diagonal \(-\tilde{\alpha}\, q^2\) penalizes large inventory; the evenness is the reflection symmetry \(q \mapsto -q\) of the generator.

Merton portfolio problem

An investor allocates a fraction \(u \in [0,1]\) of wealth \(x\) to a risky asset with drift \(\mu\) and volatility \(\sigma\), the remainder to a risk-free asset at rate \(r\), to maximize expected log terminal wealth.

Dynamics and objective

\[ \begin{aligned} \frac{dS}{S} &= \mu\,dt + \sigma\,dW, \\ dx &= x\big[r + u(\mu - r)\big]\,dt + x u \sigma\,dW, \\ J(u) &= \mathbb{E}\big[\ln x_T\big]. \end{aligned} \]

HJB equation

The value function \(V(t,x) = \sup_u J(u)\) satisfies

\[ 0 = \partial_t V + \sup_u \Big\lbrace x\big(r + u(\mu - r)\big)\partial_x V + \tfrac12 (x u \sigma)^2 \partial_x^2 V \Big\rbrace, \qquad V(T,x) = \ln x. \]

Exact solution

The log-utility ansatz \(V(t,x) = \ln x + B\,(T-t)\) is exact. The first-order condition in \(u\) gives the constant Merton fraction

\[ u^* = \frac{\mu - r}{\sigma^2}, \]

and substitution yields

\[ V(t,x) = \ln x + \left[ r + \frac{(\mu - r)^2}{2\sigma^2} \right](T - t). \]

The policy is independent of wealth and time. The wealth independence is a scale symmetry of log utility: \(x \mapsto c x\) shifts the value by the additive constant \(\ln c\), so the wealth scale drops out of the first-order condition.

Merton portfolio with deterministic jumps

A jump-diffusion extension in which the risky asset has a deterministic multiplicative jump \(y > 0\).

Dynamics

The risky asset and wealth follow

\[ \begin{aligned} \frac{dS}{S} &= \mu\,dt + \sigma\,dW + (y - 1)\,dN_t, \\ dx &= x\big[r + u(\mu - r)\big]\,dt + x u \sigma\,dW + x u (y - 1)\,dN_t, \end{aligned} \]

where \(N_t\) is a Poisson process of intensity \(\lambda\), so a jump scales wealth by \(1 + u(y-1)\). The objective remains \(J(u) = \mathbb{E}[\ln x_T]\).

HJB equation

\[ 0 = \partial_t V + \sup_u \Big\lbrace x\big[r + u(\mu-r)\big] V_x + \tfrac12 x^2 u^2 \sigma^2 V_{xx} + \lambda\big[ V\!\big(x(1 + u(y-1))\big) - V(x) \big] \Big\rbrace. \]

Exact solution

With \(V(t,x) = \ln x + B\,(T-t)\) the jump term contributes the additive constant \(\lambda \ln\big(1 + u(y-1)\big)\). The value coefficient is

\[ B = r + u^*(\mu - r) - \tfrac12 (u^*\sigma)^2 + \lambda\ln\big(1 + u^*(y-1)\big), \]

and the optimal fraction \(u^*\) solves the quadratic first-order condition

\[ 0 = (\mu - r) - \sigma^2 u^* + \lambda\,\frac{y - 1}{1 + u^*(y-1)}. \]

Setting \(\lambda = 0\) or \(y = 1\) recovers the no-jump Merton fraction \(u^* = (\mu - r)/\sigma^2\) and its value.

Merton portfolio with log-normal jumps

When the jump multiplier \(Y\) is log-normal, \(\ln Y \sim \mathcal{N}(m, \delta^2)\), the jump distribution is a continuum and the policy has no closed form.

Dynamics

\[ dx = x\big[r + u(\mu-r)\big]\,dt + x u \sigma\,dW + x u (Y - 1)\,dN_t, \]

with \(Y\) independent of the Poisson process. Log utility requires \(0 \le u \le 1\) so that wealth stays positive for any jump.

Exact solution (semi-closed form)

The ansatz \(V(t,x) = \ln x + B\,(T-t)\) remains exact, with

\[ B = r + u^*(\mu - r) - \tfrac12 (u^*\sigma)^2 + \lambda\,\mathbb{E}\big[\ln(1 + u^*(Y-1))\big], \]

and \(u^*\) solves the transcendental first-order condition

\[ 0 = (\mu - r) - \sigma^2 u^* + \lambda\,\mathbb{E}\!\left[\frac{Y - 1}{1 + u^*(Y-1)}\right]. \]

There is no closed form for \(u^*\). The expectations are evaluated by Gauss-Hermite quadrature on the standard normal \(z = (\ln Y - m)/\delta\), and \(u^*\) is bracketed on \([0,1]\) and refined by bisection. With \(\delta = 0\) the jump degenerates to the deterministic multiplier \(y = e^m\), reducing to the deterministic-jump case above.

Linear-quadratic regulator (correlated)

A finite-horizon LQ regulator with \(n\)-dimensional state and additive noise through a constant diffusion matrix \(C\).

Dynamics and objective

\[ \begin{aligned} dx &= (A x + B u)\,dt + C\,dW, \\ J(u) &= \mathbb{E}\!\left[\int_0^T \big(x^\top Q x + u^\top R u\big)\,dt + x_T^\top Q_T x_T\right], \end{aligned} \]

where the components of \(W\) are independent standard Brownian motions. The problem is to minimize \(J\), equivalently to maximize \(-J\). \(Q, Q_T\) are positive semidefinite and \(R\) is positive definite. The covariance of the noise is \(D = C C^\top\), so a full (non-diagonal) \(C\) yields correlated coordinates with \(d\langle x_i, x_j\rangle = D_{ij}\,dt\).

HJB equation

\[ 0 = \partial_t V + \sup_u \Big\lbrace -\big(x^\top Q x + u^\top R u\big) + (Ax + Bu)^\top \nabla V + \tfrac12\,\operatorname{tr}\!\big(C C^\top \operatorname{Hess} V\big) \Big\rbrace. \]

Exact solution

The value is quadratic,

\[ V(t,x) = -x^\top P(t)\,x - q(t), \qquad u^*(t,x) = -R^{-1} B^\top P(t)\,x, \]

where \(P(t)\) solves the backward Riccati ODE

\[ \dot{P} = -A^\top P - P A - Q + P B R^{-1} B^\top P, \qquad P(T) = Q_T, \]

and \(q(t)\) is the scalar, state-independent correction

\[ \dot{q} = -\operatorname{tr}\big(C C^\top P\big), \qquad q(T) = 0. \]

Because the noise is additive, \(P(t)\) and the feedback gain are independent of \(C\); the correlation enters the value only through \(\operatorname{tr}(C C^\top P)\) in \(q(t)\), and it cancels out of the optimal control. Setting \(C\) block-diagonal recovers the uncorrelated reduction.

Stationary (algebraic Riccati) reduction

In the infinite-horizon limit with discount rate \(\rho > 0\), the value is stationary and the Riccati ODE reduces to the algebraic Riccati equation. For the scalar problem \(dx = (a x + b u)\,dt + c\,dW\) with running cost \(q x^2 + r u^2\), the value is \(V(x) = -P x^2\) and the control is \(u^*(x) = -\tfrac{b}{r} P x\), where \(P\) is the positive root of

\[ \frac{b^2}{r} P^2 + (\rho - 2a) P - q = 0. \]

Linear-quadratic regulator with Poisson jumps

A scalar jump extension providing the minimal jump-diffusion control benchmark. A single Poisson source drives the jump term.

Dynamics and objective

\[ dx = (a x + b u)\,dt + \sigma\,dW + \xi\,dN_t, \]

where \(N_t\) is Poisson with intensity \(\lambda\), and at each jump the state increments by the amplitude \(\xi\), a random variable with law \(\mu\), mean \(\bar{\xi} = \mathbb{E}[\xi]\), and second moment \(\mathbb{E}[\xi^2]\). The amplitude may be deterministic. The running and terminal costs are \(q x^2 + r u^2\) and \(q_T x_T^2\) with \(q, q_T \ge 0\) and \(r > 0\); the problem maximizes the negative cost, so \(V\) below is the maximized value.

HJB equation

\[ 0 = \partial_t V + \sup_u \big\lbrace -(q x^2 + r u^2) + (a x + b u) V_x \big\rbrace + \tfrac12 \sigma^2 V_{xx} + \lambda \int \big[ V(x+\xi) - V(x) \big]\, \mu(d\xi), \]

with terminal \(V(T,x) = -q_T x^2\).

Exact solution

The ansatz \(V(t,x) = -P(t) x^2 - m(t) x - n(t)\) is exact, since the jump integral maps quadratics to quadratics. Matching powers of \(x\) gives three decoupled ODEs. The quadratic coefficient satisfies the no-jump Riccati equation

\[ \dot{P} = -2 a P - q + \frac{b^2}{r} P^2, \qquad P(T) = q_T, \]

so jumps do not enter \(P\). The linear and constant coefficients satisfy

\[ \dot{m} = -\big(a - \tfrac{b^2}{r} P(t)\big) m - 2\lambda P(t)\, \bar{\xi}, \qquad m(T) = 0, \]

\[ \dot{n} = -\sigma^2 P(t) + \frac{b^2}{4r} m(t)^2 - \lambda P(t)\, \mathbb{E}[\xi^2] - \lambda m(t)\, \bar{\xi}, \qquad n(T) = 0. \]

The optimal control keeps the linear feedback form plus an affine jump correction,

\[ u^*(t,x) = -\frac{b}{r} P(t)\, x - \frac{b}{2r} m(t). \]

For a reflection-symmetric jump law (\(\mu\) symmetric about zero, so \(\bar{\xi} = 0\)) the linear term vanishes, \(m \equiv 0\), so the control is the pure linear feedback \(u^* = -\tfrac{b}{r} P(t)\,x\) and the jump enters the value only through the constant

\[ \dot{n} = -\sigma^2 P(t) - \lambda P(t)\, \mathbb{E}[\xi^2], \qquad n(T) = 0. \]

Setting \(\lambda = 0\) recovers the scalar no-jump regulator: \(m \equiv 0\) and \(\dot n = -\sigma^2 P\) reproduce the diffusion-only correction with \(C = \sigma\). The validation targets are the sign and reduction checks: \(P(t) \ge 0\) throughout (since \(q, q_T \ge 0\)), \(V(t,x) \le 0\), the policy reduces to the no-jump feedback law when \(\lambda = 0\) or \(\bar{\xi} = 0\), and the jump contributes only a negative constant shift \(\lambda P\,\mathbb{E}[\xi^2]\) to the symmetric-case value.