If you scroll quantitative finance Twitter for any length of time you will see them: 3D surfaces that bend and warp from day to day, indexed by strike on one axis and time to expiry on the other, color-graded like a thermograph. The thing they look most like is a heat diffusion problem from a sophomore PDE class — a sharp initial feature smearing out into a Gaussian-broadened wavefront.
They are not a metaphor for heat diffusion. They are heat diffusion. The Black–Scholes PDE, under one change of variables, is the canonical 1D heat equation. Gamma is the Gaussian heat kernel up to a multiplicative factor. The vol surface you see on your screen is what you get when the market's correction to Black–Scholes' log-normality assumption is plotted as a height function. Once you see this once you cannot unsee it.
This article is about where these surfaces come from, what mathematics lives under them, and what fifty years of effort have done to the original Black–Scholes story since 1973.
Black–Scholes is the heat equation
A European option price on a non-dividend asset, under risk-neutral dynamics , satisfies
with terminal condition .
Change variables. Let , , and write for constants chosen to kill the drift and discount terms. The PDE collapses to
This is the standard heat equation on . The terminal payoff becomes the initial condition in , and the Black–Scholes formula is what you get by convolving that initial condition with the Gaussian heat kernel .
Original interactive plotRendered in-browser from the equations and parameters stated in this article.
The front edge of this surface — close to zero — is the hockey-stick payoff . As you move back in time the kink at the strike smooths out into the BS S-curve. The smoothing operator is the same Gaussian Green's function that diffuses temperature across a metal rod.
The Greeks are derivatives of a heat-equation solution
Differentiate Black–Scholes with respect to its arguments and you get the Greeks. Each one satisfies its own parabolic PDE — same operator structure, shifted coefficients, sometimes an inhomogeneous source term.
Gamma is the most striking. For a European call,
where is the standard normal density. Up to the multiplicative factor , gamma is a Gaussian. Under the log-moneyness coordinates it is exactly the heat kernel:
Original interactive plotRendered in-browser from the equations and parameters stated in this article.
This is the gamma spike that 0DTE traders post screenshots of. As time to expiry shrinks, gamma concentrates at the strike — the price becomes maximally sensitive to the underlying right where it matters. The dollar-gamma quantity , aggregated across the option-market open interest at each strike, is the famous gamma exposure (GEX) that retail microstructure FinTwit (SpotGamma, GEXStream) charts as a bar plot.
Vega satisfies an inhomogeneous Black–Scholes PDE with source term , which is why vega is concentrated near the money: vega is gamma's integral against the gamma kernel. In stochastic-vol models — where vol becomes a state variable rather than a parameter — vega is replaced by a genuine partial derivative in the variance direction, and the source-term structure folds into the bigger PDE. The classical second-order Greeks — vanna , charm , volga — are the higher-order surfaces dealers chart to predict end-of-day hedging flows.
What the volatility surface actually is
You cannot see options through the constant- lens for very long before the lens cracks. The model assumes one volatility; the market quotes a different price for every strike and every maturity. Invert the Black–Scholes formula option by option and you get the implied volatility surface — the that, plugged into BS, reproduces each market mid-price.
If Black–Scholes were correct the surface would be flat. It is not. Its shape encodes precisely what the model misses: fat tails, downside skew, jump risk, vol-of-vol, the inability of a single Gaussian to do all the work.
Original interactive plotRendered in-browser from the equations and parameters stated in this article.
The picture has a name. SPX surfaces always have a downward slope in strike — out-of-the-money puts cost more than out-of-the-money calls, the smirk that has been the equity-vol signature since 1987. FX surfaces are symmetric smiles. The slope encodes the implied probability of a tail event. The curvature encodes the implied kurtosis. The whole surface flexes daily as the market revises its tail estimates.
The standard arbitrage-free parameterization is Gatheral's SVI (Madrid 2004, arbitrage-free conditions in Gatheral and Jacquier 2014, arXiv:1204.0646). For each maturity, the total implied variance as a function of log-moneyness takes the form
Five parameters per slice, asymptotically linear in both wings (Roger Lee's moment formula), arbitrage-free under explicit Durrleman conditions on the second derivative. The surface SVI (SSVI) extension imposes calendar- and butterfly-arbitrage constraints globally and is what most modern fitters use.
Dupire's forward equation
The local volatility model takes the surface as an input and asks for the unique deterministic function such that the diffusion reproduces every observed European option price. Dupire (1994) gave the closed-form inverse:
Mechanically this is the inverse of a forward parabolic PDE — the Fokker–Planck equation for the risk-neutral density of , integrated against the call payoff and differentiated twice in strike. Same parabolic structure as the BS PDE, axes flipped: instead of "given , find ", you get "given , find ".
Local vol fits today's surface exactly, by construction. But it predicts forward smiles that flatten too quickly. A barrier option or a forward-starting option priced under local vol typically misprices by a percentage point of vol relative to the same option priced under stochastic vol, even when both models agree on every vanilla price. Hagan, Kumar, Lesniewski, Woodward ("Managing Smile Risk", 2002) made the critique sharp: local vol moves the smile in the wrong direction when spot moves. Exotic desks generally use a local-stochastic hybrid (LSV — see Guyon and Henry-Labordère 2012) that calibrates exactly to vanillas like local vol but produces realistic forward-smile dynamics like stochastic vol.
Stochastic volatility: now you have a 2D PDE
To get realistic forward smiles you let volatility itself be random. The canonical model is Heston (1993):
The variance is a CIR (square-root) process, mean-reverting to at rate , with vol-of-vol . The Feller condition keeps almost surely. The pricing PDE is now 2-dimensional in :
Heston's contribution was making this tractable. The characteristic function of has the affine form
with satisfying a system of Riccati ODEs in . European prices follow from inverting the characteristic function by FFT — the Carr–Madan (1999) trick that made stochastic-vol calibration competitive with closed-form BS.
The correlation for equity is what produces skew: when spot falls, vol rises, OTM puts get bid. The vol-of-vol controls smile curvature. Heston gets the gross shape right but cannot fit the short-end skew of SPX without unreasonable parameters — which is the empirical observation that motivates rough volatility below.
The interest-rate analogue is SABR (Hagan, Kumar, Lesniewski, Woodward 2002):
A singular-perturbation expansion gives a closed-form approximation for that has been the swaption desk's working model for two decades. Adding jumps to Heston gives Bates (1996): an exponential-Lévy jump term, retaining the affine structure, sacrificing tractability of the time-stepping but keeping FFT pricing.
Jumps: parabolic plus a non-local operator
A continuous Brownian motion cannot produce a 5-sigma Monday. Merton (1976) added log-normal jumps:
with a Poisson process with intensity and jump size . The pricing operator picks up an integral term:
This is a partial integro-differential equation (PIDE). The local Black–Scholes part smooths; the non-local jump kernel kicks the surface around. Kou (2002) replaces Gaussian jumps with an asymmetric double-exponential distribution, giving closed-form barrier and lookback prices. The CGMY model (Carr, Geman, Madan, Yor 2002) is a pure-jump Lévy process with density
where the four parameters tune activity, asymmetric tail decay, and the fine-structure index . CGMY is the workhorse pure-jump model in equity options, used whenever the tail behavior matters more than the diffusion.
Numerically, PIDEs are solved by splitting the local operator (treated implicitly via finite differences) from the non-local jump integral (treated explicitly via FFT). Cont and Voltchkova (2005) is the standard reference.
Rough volatility
The cleanest single-line empirical paper in mathematical finance is Gatheral, Jaisson, and Rosenbaum, "Volatility is rough" (arXiv:1410.3394, 2014). Take a high-frequency time series of realized variance, measure the scaling of its increments, fit the Hurst parameter . The Brownian baseline is . The empirical value, across asset classes and time periods, is .
Realized volatility is rougher than Brownian motion. The fractional Brownian motion that drives it has paths with Hölder regularity 0.1, not 0.5 — visibly more jagged at every scale.
Bayer, Friz, and Gatheral built the rough Bergomi model on this empirical regularity (arXiv:1502.02944, 2016):
The kernel is singular at the origin when . The variance process is no longer Markovian in , and consequently no finite-dimensional PDE exists for the option price. Pricing falls back to Monte Carlo or to specialized series expansions.
Under the rough Heston model with Hurst parameter , the characteristic function of satisfies
where and solves the fractional Riccati equation
with and the fractional Caputo derivative. As this reduces to the classical Heston Riccati system.
The fractional Riccati replaces the ODE Riccati of classical Heston, and the price of admission is one nonlocal integral operator per time step. The same FFT-after-Riccati pipeline still works; it just costs more.
A different way to recover tractability is to lift the rough model to a higher-dimensional Markovian one. Abi Jaber (arXiv:1810.04868, 2019) and Abi Jaber–El Euch (arXiv:1801.10359, 2019) showed that the fractional kernel admits a Laplace representation
and truncating the integral to a finite sum of exponentials gives an -dim Markovian approximation that converges to rough Heston as . PDE methods come back.
The single sharpest argument for rough volatility being the right framework is short-end skew scaling. Empirically, the slope of the SPX implied vol smile at the money goes like as — i.e. it blows up like . Every Markovian stochastic vol model (Heston, SABR, anything with ) predicts a finite-slope short-end skew. Rough vol with fits the scaling. No model with does.
The empirical status of rough vol is not entirely settled. Cont and Das (arXiv:2203.13820, 2022) argue that observed roughness may be a microstructure-noise artifact and that the true may be closer to . Mouti (arXiv:2312.01426) finds on range-based proxies. The SIAM volume Rough Volatility (Bayer–Friz–Fukasawa–Gatheral–Jacquier–Rosenbaum, 2023) is the consolidated reference; this is an open empirical debate as much as a theoretical one.
High dimensions: the BSDE-as-neural-network move
Suppose you need to price a basket option on underlying assets. The pricing PDE lives in 50 spatial dimensions plus time. A finite-difference grid with points per dimension needs memory cells. This is the curse of dimensionality: PDE methods are dead beyond , and until 2017 the only option was Monte Carlo with all its variance-reduction machinery.
E, Han, and Jentzen changed this. Their deep BSDE solver (arXiv:1707.02568, PNAS 2018) reformulates the semilinear parabolic PDE
as a forward-backward SDE
where and . Parameterize the gradient as a deep neural network at each time step. Simulate forward paths of . Train the network by minimizing — that is, match the predicted terminal value to the actual payoff in expectation. The result is a learned function that approximates the PDE solution in 100 dimensions in minutes.
Sirignano and Spiliopoulos took the more direct route. Their Deep Galerkin Method (arXiv:1708.07469, J. Comp. Phys. 2018) parameterizes as a deep net directly, samples uniformly from the domain, and minimizes
Mesh-free, tested up to 200 dimensions on free-boundary problems and HJB equations. The general principle that came out of these two papers — that neural-network parameterization of the solution beats grid methods in high dimensions, with a complexity polynomial in rather than exponential — has been formally proven for semilinear Kolmogorov PDEs (Hutzenthaler, Jentzen, Kruse, 2020).
For mathematical finance, the practical effect is that 50- and 100-dimensional basket options, XVA valuation adjustments, and credit derivative pricing — all of which were previously locked into expensive Monte Carlo pipelines — now have neural PDE solvers that run in minutes on commodity GPUs.
What the 2026 frontier looks like
Three threads dominate current research.
Path-dependent volatility. Guyon and Lekeufack (SSRN 4174589, 2022; Quantitative Finance 2023) claim — controversially — that 90% of the variance of future implied volatility is explained deterministically by past returns and past squared returns, with no stochastic component needed. Their 4-factor PDV model uses two-exponential-kernel mixtures
with . Markovian, low-parametric, fits both SPX and VIX smiles jointly — a feat that has eluded stochastic and rough vol models for two decades. Gazzani and Guyon (arXiv:2406.02319) extended this to a full pricing and calibration framework in 2024.
Joint SPX/VIX calibration. The "holy grail" of equity vol modeling: one model fitting both surfaces simultaneously. The 2023–2025 entrants:
- Signature-based models (Cuchiero, Gazzani, Möller, Svaluto-Ferro, arXiv:2301.13235, 2023). Make volatility a linear functional of the path signature — the universal feature map from rough path theory. The VIX squared then becomes a polynomial in the signature components, computable in closed form.
- Gaussian polynomial volatility (Abi Jaber, Illand, Li, arXiv:2212.08297). Volatility is a polynomial of a Gaussian Volterra process. The empirically-winning version is, surprisingly, not rough — it is a one-factor Markovian Gaussian model. The rough-vol case is included as a special case but not preferred.
Deep hedging. Buehler, Gonon, Teichmann, and Wood (arXiv:1802.03042, Quantitative Finance 2019) reformulated hedging itself as a reinforcement learning problem. Train a neural-network policy to minimize a convex risk measure (CVaR, entropic risk) of terminal P&L, including transaction costs, market impact, and liquidity constraints. None of those fit into the classical PDE framework. The trained policy approximates the optimal hedge in a model-free way.
This is the single biggest practical innovation in derivatives pricing since Heston, and it is in production at multiple banks. The mathematical structure underneath is a stochastic optimal control problem whose value function satisfies an HJB equation — but the HJB is in too many dimensions to solve by grid methods, so you let a neural net approximate the policy directly. The same trick as deep BSDE, applied to control rather than valuation.
How the FinTwit surfaces map to all of this
| Visualization | Underlying object | Generated by |
|---|---|---|
| Implied vol surface | The market's correction to BS log-normality | Inverting BS option by option; fitted by SVI/SSVI |
| Local vol surface | The unique Markovian diffusion reproducing the vanilla surface | Dupire's forward equation |
| Gamma / vanna / charm surface | Sensitivities of the BS price to spot, vol, time | Differentiation of the BS formula; same parabolic operator |
| GEX-by-strike bar plot | Dealer net dollar-gamma at each strike | Aggregate of across the option chain |
| Heat-diffusion movie | The BS PDE under change of variables | , literally |
| Heston / rough Heston smile fit | Calibration residuals for stochastic vol models | The 2D Heston PDE / fractional Riccati |
| Rough vs Brownian path comparison | Sample paths of fractional Brownian motion vs standard BM | Volterra integral with singular kernel |
| VIX term structure | Forward variance curve | Variance swap arithmetic |
| Monte Carlo fan chart | Sample paths of an SDE | GBM / Heston / jump-diffusion / rough Bergomi simulation |
| Yield-curve evolution surface | Forward rate over (calendar time, maturity) | HJM forward-rate dynamics |
Most of these are parabolic-PDE solutions or moments thereof. The "wave equation" feel that animated surfaces sometimes have is misleading — they are not hyperbolic. Each animation frame is a separately-calibrated diffusion, and the apparent wave is just the market's day-to-day revision of its risk premiums.
The practical question
Every model here is correct for the regime it was built for and wrong for one it was not. A liquid vanilla desk calibrates SVI or Heston daily and worries about delta and gamma. An exotic desk needs forward-smile dynamics that local vol misses. A multi-asset book on basket exotics needs deep BSDE or DGM. A market-making operation with transaction costs needs deep hedging. A short-end SPX trader needs rough vol to get the skew right.
The practical question is always: which model is wrong in the way that matters least for the trade you are pricing? That has no generic answer. It is a judgment call, made by someone who knows what each model gets right and what it gets wrong, and it is what separates good derivatives desks from bad ones.
The mathematics is one continuous story. Parabolic PDEs, sometimes coupled, sometimes integro-differential, sometimes fractional, sometimes solved by neural networks instead of grids. The story has been running for fifty-three years since Black and Scholes (1973), and the equations keep accommodating new pieces of empirical reality without ever quite invalidating the previous generation. The implied vol surface is still the most important object in equity derivatives. Black–Scholes is still the formula every option price is quoted in.
The 3D surfaces on quant Twitter are the visible surface of all of this. They are not metaphors. They are heat equations.
Comments