Pathwise derivatives.
Local control recovery.
We study stochastic control in financial mathematics, with a central focus on local recovery using pathwise derivatives. We ask how information obtained from simulated trajectories can determine a control at a given time and state, and how these recovery steps can be organized into policy-update operators.
Our research spans regular control, delay, transaction costs, partial observation, and deep hedging. Across these problems, we investigate where the recovery principle applies, where it reaches its limits, and what modifications make further extensions possible.
Local recovery
A common starting point is to simulate trajectories under a candidate policy and differentiate the relevant path functional. The resulting information enters a local control problem: maximize the appropriate Hamiltonian over admissible controls.
- 01Differentiate pathsExtract the information needed for control.
- 02Maximize the HamiltonianUse the PMP or HJB Hamiltonian.
- 03Recover a controlChoose an admissible action at the current state.
Schematic regular-control form
Here, \(\mathcal I^{\pi}\) denotes the information extracted under the candidate policy, and \(\mathcal U\) is the admissible control set. The information and Hamiltonian depend on the formulation.
The maximization may be solved analytically or numerically. With constraints, KKT conditions can help characterize candidate solutions; identifying a maximizer requires the appropriate conditions. A control recovered from a candidate policy is then assessed for feasibility, accuracy, and policy improvement.
Two derivative routes
OL-BPTT · Open-loop differentiation
Adjoint information → PMP recoveryWe use open-loop backpropagation through time to estimate the adjoint quantities associated with the control problem’s BSDE. The recovered adjoint quantities enter a Pontryagin maximum principle (PMP) Hamiltonian for local control recovery.
CL-BPTT · Closed-loop differentiation
Value derivatives → HJB recoveryWith policy parameters held fixed, we differentiate and average rollouts while retaining the state-to-policy feedback dependence. This estimates local derivatives of the candidate policy's value, which enter a Hamilton–Jacobi–Bellman (HJB) Hamiltonian for local control recovery.
These routes use different differentiation conventions. Relating their information and their induced control updates is part of our research.
Which derivative information is needed?
The answer depends on the problem. The HJB Hamiltonian can require second-order value derivatives when the diffusion depends on the control. On the PMP side, the required adjoint system also depends on the formulation. We study these distinctions rather than treating every pathwise derivative as the same quantity.
Recovery as an operator
Local recovery defines a policy update when it can be assembled into an admissible policy. We view information extraction and control recovery as two components of an operator:
\(\mathcal E\) extracts the required information; \(\mathcal R\) performs local recovery.
This perspective connects one-step recovery with iterative policy refinement. HJB-based policy iteration provides a reference point. A central question is whether, and under what conditions, an OL-BPTT/PMP-based operator also has a policy-iteration structure.
- Policy improvement
- When does a recovery step improve the objective?
- Fixed points
- What optimality conditions does an unchanged policy satisfy?
- Convergence
- When do repeated updates converge, and to what?
- Decision accuracy
- How do approximation and derivative errors affect the recovered control?
We examine these properties for each problem class. Defining an iteration is a starting point; establishing improvement, convergence, or optimality requires further analysis.
Extensions and limits
The breadth of our research comes from testing a common principle against different control structures. When direct recovery is insufficient, we ask what additional state, information, optimality conditions, or update mechanisms are needed.
Portfolio and consumption decisions, state and control constraints, and control-dependent dynamics.
History-dependent states and costs, with corresponding changes to differentiation and adjoint information.
No-trade regions and trading decisions that may require gradient constraints, variational inequalities, or intervention conditions.
Control with filtering information: how the information state changes the recovery problem and its computation.
Dynamic hedging policies under market frictions and constraints, connecting learned decisions with control structure.
Our aim is to identify both useful extensions and genuine boundaries. Understanding why a recovery rule fails, and how it must be modified, is part of the research program.
Selected research
Methods such as PG-DPO and BG-DPO are concrete developments within this program. The following papers and ongoing projects illustrate its different directions.
- Adjoint-to-control recoveryScalable Pontryagin-Guided Adjoint-to-Control Recovery for Constrained Dynamic Portfolio Choice ↗
- Policy-update operatorsSelf-Consistent Adjoint Policy Iteration for Constrained Dynamic Portfolio Choice ↗
- Delay and path dependenceDifferentiating Through Delay: Stochastic Control Without Full History Hessians
- Local recovery with transaction costsKnowing When Not to Act: Latent No-Action Region Recovery Hidden in Neural Control
- Deep hedging · Work in progressDeep Hedging through Bellman-Guided Direct Policy Optimization
Earlier technical material
BPTT as a Pathwise Costate Solver develops the earlier PG-DPO perspective.