MLFE LAB

Research

Stochastic control through pathwise differentiation and local recovery.

Pathwise derivatives.
Local control recovery.

We study stochastic control in financial mathematics, with a central focus on local recovery using pathwise derivatives. We ask how information obtained from simulated trajectories can determine a control at a given time and state, and how these recovery steps can be organized into policy-update operators.

Our research spans regular control, delay, transaction costs, partial observation, and deep hedging. Across these problems, we investigate where the recovery principle applies, where it reaches its limits, and what modifications make further extensions possible.

Local recovery

A common starting point is to simulate trajectories under a candidate policy and differentiate the relevant path functional. The resulting information enters a local control problem: maximize the appropriate Hamiltonian over admissible controls.

  1. 01Differentiate pathsExtract the information needed for control.
  2. 02Maximize the HamiltonianUse the PMP or HJB Hamiltonian.
  3. 03Recover a controlChoose an admissible action at the current state.

Schematic regular-control form

\[u^{+}(t,x)\in\operatorname*{arg\,max}_{u\in\mathcal U(t,x)}\;\mathcal H\!\left(t,x,u;\mathcal I^{\pi}(t,x)\right).\]
\[\begin{gathered}u^{+}(t,x)\in\\[4pt]\operatorname*{arg\,max}_{u\in\mathcal U(t,x)}\;\mathcal H\!\left(t,x,u;\mathcal I^{\pi}(t,x)\right).\end{gathered}\]

Here, \(\mathcal I^{\pi}\) denotes the information extracted under the candidate policy, and \(\mathcal U\) is the admissible control set. The information and Hamiltonian depend on the formulation.

The maximization may be solved analytically or numerically. With constraints, KKT conditions can help characterize candidate solutions; identifying a maximizer requires the appropriate conditions. A control recovered from a candidate policy is then assessed for feasibility, accuracy, and policy improvement.

Two derivative routes

OL-BPTT · Open-loop differentiation

Adjoint information → PMP recovery

We use open-loop backpropagation through time to estimate the adjoint quantities associated with the control problem’s BSDE. The recovered adjoint quantities enter a Pontryagin maximum principle (PMP) Hamiltonian for local control recovery.

CL-BPTT · Closed-loop differentiation

Value derivatives → HJB recovery

With policy parameters held fixed, we differentiate and average rollouts while retaining the state-to-policy feedback dependence. This estimates local derivatives of the candidate policy's value, which enter a Hamilton–Jacobi–Bellman (HJB) Hamiltonian for local control recovery.

These routes use different differentiation conventions. Relating their information and their induced control updates is part of our research.

Which derivative information is needed?

The answer depends on the problem. The HJB Hamiltonian can require second-order value derivatives when the diffusion depends on the control. On the PMP side, the required adjoint system also depends on the formulation. We study these distinctions rather than treating every pathwise derivative as the same quantity.

Recovery as an operator

Local recovery defines a policy update when it can be assembled into an admissible policy. We view information extraction and control recovery as two components of an operator:

\[\mathcal T=\mathcal R\circ\mathcal E,\qquad\pi_{k+1}=\mathcal T(\pi_k).\]

\(\mathcal E\) extracts the required information; \(\mathcal R\) performs local recovery.

This perspective connects one-step recovery with iterative policy refinement. HJB-based policy iteration provides a reference point. A central question is whether, and under what conditions, an OL-BPTT/PMP-based operator also has a policy-iteration structure.

Policy improvement
When does a recovery step improve the objective?
Fixed points
What optimality conditions does an unchanged policy satisfy?
Convergence
When do repeated updates converge, and to what?
Decision accuracy
How do approximation and derivative errors affect the recovered control?

We examine these properties for each problem class. Defining an iteration is a starting point; establishing improvement, convergence, or optimality requires further analysis.

Extensions and limits

The breadth of our research comes from testing a common principle against different control structures. When direct recovery is insufficient, we ask what additional state, information, optimality conditions, or update mechanisms are needed.

Regular stochastic control

Portfolio and consumption decisions, state and control constraints, and control-dependent dynamics.

Delay and path dependence

History-dependent states and costs, with corresponding changes to differentiation and adjoint information.

Transaction costs

No-trade regions and trading decisions that may require gradient constraints, variational inequalities, or intervention conditions.

Partial observation and Zakai dynamics

Control with filtering information: how the information state changes the recovery problem and its computation.

Deep hedging

Dynamic hedging policies under market frictions and constraints, connecting learned decisions with control structure.

Our aim is to identify both useful extensions and genuine boundaries. Understanding why a recovery rule fails, and how it must be modified, is part of the research program.

Selected research

Methods such as PG-DPO and BG-DPO are concrete developments within this program. The following papers and ongoing projects illustrate its different directions.

Earlier technical material

BPTT as a Pathwise Costate Solver develops the earlier PG-DPO perspective.