Koopman operator: adjoint, normality, SVD

· Revised Jun 01, 2026

TL;DR. What is the adjoint of the Koopman operator on a stochastic Markov process, when is the operator normal, and why does the singular value decomposition become the natural object the moment normality fails?

The Koopman operator turns nonlinear dynamics into linear analysis at the price of an infinite-dimensional function space. Three questions about it come up repeatedly when people work in this language: what is the adjoint, when is the operator normal, and why does the singular value decomposition rather than the eigenvalue decomposition show up as the natural object in most modern methods? This post answers each in turn, in a way that resolves a confusion that comes up often: that the adjoint of the Koopman operator is its reverse-time version. It is not, except in a special case.

Setup: the Koopman operator on a Markov process

Let (X,F)(\mathcal{X}, \mathcal{F}) be a measurable state space and consider a time-homogeneous Markov process with transition density p(xx)p(\mathbf{x}' \mid \mathbf{x}). For an observable g ⁣:XRg \colon \mathcal{X} \to \mathbb{R}, define the Koopman operator

(Kg)(x)Ep(xx)[g(x)]=g(x)p(xx)dx.(\mathcal{K} g)(\mathbf{x}) \triangleq \mathbb{E}_{p(\mathbf{x}' \mid \mathbf{x})}[g(\mathbf{x}')] = \int g(\mathbf{x}') \, p(\mathbf{x}' \mid \mathbf{x}) \, d\mathbf{x}'.

In the deterministic case xt+1=T(xt)\mathbf{x}_{t+1} = T(\mathbf{x}_t) this reduces to composition, Kg=gT\mathcal{K} g = g \circ T. In the stochastic case it is the conditional expectation operator. Either way, K\mathcal{K} pushes observables forward in time: given the value of gg at the future state, Kg\mathcal{K} g is its conditional expectation given the present.

Two distributions matter. Let ρ0\rho_0 be the marginal of the present state x\mathbf{x} and ρ1\rho_1 the marginal of the future state x\mathbf{x}'. We work in Hilbert spaces of square-integrable observables, L2(ρ0)L^2(\rho_0) and L2(ρ1)L^2(\rho_1), with the standard inner products f,gρ=fgdρ\langle f, g \rangle_\rho = \int f g \, d\rho. The Koopman operator K\mathcal{K} then maps L2(ρ1)L2(ρ0)L^2(\rho_1) \to L^2(\rho_0): forward in time, but backward in the direction of the marginals. It takes an observable of the future state and returns one of the present state.

Already at this stage the setup is more delicate than the usual “Koopman acts on functions” presentation: K\mathcal{K} is an operator between two different Hilbert spaces unless something special makes ρ0=ρ1\rho_0 = \rho_1.

The adjoint is the backward predictor

What is K\mathcal{K}^*? By the defining identity of the adjoint,

Kg,fρ0=g,Kfρ1\langle \mathcal{K} g, f \rangle_{\rho_0} = \langle g, \mathcal{K}^* f \rangle_{\rho_1}

for all fL2(ρ0)f \in L^2(\rho_0), gL2(ρ1)g \in L^2(\rho_1). Expanding the left side,

Kg,fρ0=f(x)Ep(xx)[g(x)]ρ0(dx)=f(x)g(x)p(xx)ρ0(dx)dx.\langle \mathcal{K} g, f \rangle_{\rho_0} = \int f(\mathbf{x}) \, \mathbb{E}_{p(\mathbf{x}' \mid \mathbf{x})}[g(\mathbf{x}')] \, \rho_0(d\mathbf{x}) = \iint f(\mathbf{x}) g(\mathbf{x}') p(\mathbf{x}' \mid \mathbf{x}) \rho_0(d\mathbf{x}) d\mathbf{x}'.

For this to equal g(x)(Kf)(x)ρ1(dx)\int g(\mathbf{x}') (\mathcal{K}^* f)(\mathbf{x}') \rho_1(d\mathbf{x}') for every gg,

(Kf)(x)=f(x)ρ0(x)p(xx)ρ1(x)dx=Eq(xx)[f(x)](\mathcal{K}^* f)(\mathbf{x}') = \int f(\mathbf{x}) \frac{\rho_0(\mathbf{x}) \, p(\mathbf{x}' \mid \mathbf{x})}{\rho_1(\mathbf{x}')} \, d\mathbf{x} = \mathbb{E}_{q(\mathbf{x} \mid \mathbf{x}')}[f(\mathbf{x})]

where q(xx)ρ0(x)p(xx)/ρ1(x)q(\mathbf{x} \mid \mathbf{x}') \triangleq \rho_0(\mathbf{x}) p(\mathbf{x}' \mid \mathbf{x}) / \rho_1(\mathbf{x}') is the Bayes-flipped conditional: the conditional density of the past given the future.

The adjoint of the Koopman operator is the backward predictor, the conditional expectation of past given future. This is also exactly the Perron–Frobenius (transfer) operator, viewed in the correct Hilbert space pair: K\mathcal{K}^* pulls densities backward in time, while K\mathcal{K} pushes observables forward.

Remark. A common shorthand identifies K\mathcal{K}^* with the reverse-time Koopman operator ggT1g \mapsto g \circ T^{-1}, but the two coincide only when TT is invertible and measure-preserving (so q(xx)q(\mathbf{x} \mid \mathbf{x}') collapses to a point mass at T1(x)T^{-1}(\mathbf{x}')); outside that case, K\mathcal{K}^* is a genuine integral against the Bayes-flipped conditional, not a composition.

Normality

The Koopman operator is normal if KK=KK\mathcal{K} \mathcal{K}^* = \mathcal{K}^* \mathcal{K}. For this equality to even make sense, both sides must act on the same Hilbert space, which forces ρ0=ρ1=π\rho_0 = \rho_1 = \pi. Normality is therefore a property defined only in the stationary regime, and we assume K ⁣:L2(π)L2(π)\mathcal{K} \colon L^2(\pi) \to L^2(\pi) for the rest of this section. Geometrically, normality then means K\mathcal{K} and K\mathcal{K}^* share eigenspaces and L2(π)L^2(\pi) splits into mutually orthogonal K\mathcal{K}-invariant subspaces (the spectral theorem); algebraically, it is the condition that lets us diagonalize with an orthonormal basis.

Dynamically, KK\mathcal{K}^* \mathcal{K} has a single clean reading on L2(π)L^2(\pi): it is the one-step predictability operator. Its quadratic form is the squared L2L^2 norm of the forward conditional expectation,

KKg,gπ=Eπ ⁣[E[g(xt+1)xt]2]=Kgπ20,\langle \mathcal{K}^* \mathcal{K} g,\, g \rangle_\pi \,=\, \mathbb{E}_\pi\!\bigl[\, \mathbb{E}[g(\mathbf{x}_{t+1}) \mid \mathbf{x}_t]^2 \,\bigr] \,=\, \| \mathcal{K} g \|_\pi^2 \,\ge\, 0,

so KK\mathcal{K}^* \mathcal{K} is positive self-adjoint as a Gram operator (its quadratic form is a norm-squared, hence non-negative by construction).

The law of total variance pins down the spectrum precisely. For gg with Eπ[g]=0\mathbb{E}_\pi[g] = 0 and gπ=1\|g\|_\pi = 1,

1=Var(g(xt+1))=E[Var(g(xt+1)xt)]unpredictable noise+KKg,gπexplained,1 \,=\, \mathrm{Var}\bigl(g(\mathbf{x}_{t+1})\bigr) \,=\, \underbrace{\mathbb{E}\bigl[\mathrm{Var}(g(\mathbf{x}_{t+1}) \mid \mathbf{x}_t)\bigr]}_{\text{unpredictable noise}} \,+\, \underbrace{\langle \mathcal{K}^* \mathcal{K} g, g\rangle_\pi}_{\text{explained}},

so KKg,gπ[0,1]\langle \mathcal{K}^* \mathcal{K} g, g\rangle_\pi \in [0, 1] is exactly the one-step forecast R2R^2 of the observable. Plugging the ii-th right singular function ψi\psi_i, σi2=KKψi,ψiπ[0,1]\sigma_i^2 = \langle \mathcal{K}^* \mathcal{K} \psi_i, \psi_i\rangle_\pi \in [0, 1] is the ii-th-best one-step-forecast R2R^2 in the dynamics. Values of σi\sigma_i near 11 mark persistent observables; values near 00 mark observables that one transition decorrelates.

Normality, KK=KK\mathcal{K}^* \mathcal{K} = \mathcal{K} \mathcal{K}^*, equates the forward predictability operator with its time-reversed counterpart; equivalently Kgπ=Kgπ\| \mathcal{K} g \|_\pi = \| \mathcal{K}^* g \|_\pi, so forward and backward predictions of any gg have the same L2L^2 norm.

When is the Koopman operator normal? Three regimes worth distinguishing.

Measure-preserving invertible dynamics. Hamiltonian flows, ergodic group actions, volume-preserving diffeomorphisms. Here K\mathcal{K} is an isometry on L2(μ)L^2(\mu): composition with a measure-preserving invertible map preserves the inner product. In fact K\mathcal{K} is unitary, KK=KK=I\mathcal{K}^* \mathcal{K} = \mathcal{K} \mathcal{K}^* = I, so every σi=1\sigma_i = 1: invertibility loses no information about g(xt+1)g(\mathbf{x}_{t+1}) given xt\mathbf{x}_t, and every observable is fully forecast by one step. Unitary operators are normal, the spectrum lives on the unit circle, and eigenvalues are pure phases. This is the classical setting introduced by Koopman (1931).

Reversible Markov processes. A Markov process is time-reversible if and only if it satisfies the detailed balance condition π(x)p(xx)=π(x)p(xx)\pi(\mathbf{x}) p(\mathbf{x}' \mid \mathbf{x}) = \pi(\mathbf{x}') p(\mathbf{x} \mid \mathbf{x}'), where π\pi is the stationary distribution. Substituting into the backward-predictor formula, q(xx)=p(xx)q(\mathbf{x} \mid \mathbf{x}') = p(\mathbf{x} \mid \mathbf{x}'), so K=K\mathcal{K}^* = \mathcal{K}: the operator is self-adjoint on L2(π)L^2(\pi). The predictability operator then coincides with two-step forward propagation, KK=K2\mathcal{K}^* \mathcal{K} = \mathcal{K}^2, because detailed balance makes forward and backward transitions statistically indistinguishable. Eigenvalues are real, eigenfunctions are orthonormal in L2(π)L^2(\pi), and the SVD coincides with the eigenvalue decomposition with σi=λi\sigma_i = |\lambda_i|. Detailed balance is restrictive (equilibrium statistical mechanics yes, generic dissipative dynamics no), but it is the case for which the operator-theoretic story is cleanest.

Everything else. Non-invertible deterministic dynamics (the doubling map x2xmod1x \mapsto 2x \bmod 1 is the textbook example): K\mathcal{K} is no longer isometric. Composition with a many-to-one map collapses information about which preimage was used, and the squared norm of Kg\mathcal{K} g is generally less than that of gg. KKKK\mathcal{K} \mathcal{K}^* \ne \mathcal{K}^* \mathcal{K}. Stochastic Markov processes with ρ0ρ1\rho_0 \ne \rho_1 (non-stationary trajectory): even the operator’s domain and codomain are different Hilbert spaces, so normality fails for a basic type-theoretic reason. Stationary but irreversible processes (no detailed balance): K\mathcal{K} and K\mathcal{K}^* both act on L2(π)L^2(\pi) but commute only on specific subspaces. In every sub-case, KKKK\mathcal{K}^* \mathcal{K} \ne \mathcal{K} \mathcal{K}^*: forward and backward predictability define distinct singular subspaces, and the operator carries an irreducible directional asymmetry that no eigenbasis of K\mathcal{K} alone can absorb.

The summary is short. Most physically interesting dynamics give non-normal Koopman operators. Irreversible chemical kinetics, dissipative fluid flow, biased Langevin dynamics, neural-network dynamics under SGD: all non-normal. Normal Koopman is the exception.

SVD is the right tool when normality fails

When K\mathcal{K} is normal, the eigenvalue decomposition gives an orthonormal basis of eigenfunctions and the spectrum has clean physical meaning: phases on the unit circle for unitary, decay rates on the real line for self-adjoint. Eigenfunctions are Koopman modes in the classical sense, and methods like DMD that estimate eigenpairs are entirely natural.

When K\mathcal{K} is non-normal, two things break.

Eigenfunctions need not be orthogonal, and may not form a basis. For a non-normal operator the eigenfunctions can be skewed relative to each other. Even when they span the space, decomposing an arbitrary observable in the eigenbasis can require huge expansion coefficients with massive cancellation, because the basis is poorly conditioned. The eigenvalues, considered in isolation, can also be misleading: pseudospectrum-style arguments show that for non-normal operators, small perturbations to K\mathcal{K} can move eigenvalues by amounts unrelated to the perturbation magnitude. Stability conclusions drawn from spectral data alone become unreliable.

Numerical eigendecomposition becomes unstable. The gradient of an eigendecomposition with respect to the underlying matrix scales inversely with eigenvalue gaps: λi/M1/(λiλj)\partial \lambda_i / \partial M \sim 1 / (\lambda_i - \lambda_j) for close eigenvalues. Empirical second-moment matrices estimated from finite samples are perturbations of the true ones, and when the true spectrum has clustered or nearly-coincident eigenvalues, the empirical eigenvectors are not even consistent estimators of the truth, let alone differentiable in a way useful for training a neural network. This is why frameworks that backpropagate through eigh or matrix square roots (VAMPnet (Mardt et al., 2018) and DPNet (Kostic et al., 2024), among others) are hard to scale to high-dimensional, weakly-stationary, ill-conditioned regimes.

The singular value decomposition survives both. Write

K=i1σiϕiψi\mathcal{K} = \sum_{i \ge 1} \sigma_i \, \phi_i \otimes \psi_i

where {ϕi}L2(ρ0)\{\phi_i\} \subset L^2(\rho_0) and {ψi}L2(ρ1)\{\psi_i\} \subset L^2(\rho_1) are orthonormal bases of the respective spaces, and σ1σ20\sigma_1 \ge \sigma_2 \ge \cdots \ge 0. The singular values quantify how much of each input direction survives the forward operator:

σi2=KKψi,ψiρ1.\sigma_i^2 = \langle \mathcal{K}^* \mathcal{K} \, \psi_i, \, \psi_i \rangle_{\rho_1}.

That is, σi2\sigma_i^2 is the eigenvalue of the self-adjoint operator KK\mathcal{K}^* \mathcal{K} on its iith eigenfunction. The right singular functions ψi\psi_i are the directions of best forward predictability; the left singular functions ϕi\phi_i are where they land. Three facts make SVD the right primitive for non-normal K\mathcal{K}:

  • It exists for any compact operator regardless of normality.
  • The orthonormality of {ϕi}\{\phi_i\} and {ψi}\{\psi_i\} follows from the self-adjointness of KK\mathcal{K}^* \mathcal{K} and KK\mathcal{K} \mathcal{K}^*, not from any assumption on K\mathcal{K} itself.
  • The truncated SVD i=1kσiϕiψi\sum_{i=1}^k \sigma_i \phi_i \otimes \psi_i is the optimal rank-kk approximation of K\mathcal{K} in Hilbert–Schmidt norm. This is the Eckart–Young theorem (Eckart and Young, 1936) for operators, and it holds without any normality assumption.

The reversible / self-adjoint case is the degenerate special case in which SVD and eigendecomposition coincide: σi=λi\sigma_i = |\lambda_i| and ϕi=ψi\phi_i = \psi_i up to sign. So: eigendecomposition is what we can do when normality lets us; SVD is what we can always do.

When the SVD exists: compactness via stochastic regularization

The decomposition K=iσiϕiψi\mathcal{K} = \sum_i \sigma_i \phi_i \otimes \psi_i quietly assumed K\mathcal{K} is compact, so that σn0\sigma_n \to 0 and the series converges. Compactness is not automatic, and the two basic regimes split cleanly.

Deterministic measure-preserving dynamics. As established under normality, an invertible measure-preserving TT makes K\mathcal{K} a unitary isometry with every σi=1\sigma_i = 1; it is therefore not compact. The spectrum lives continuously on the unit circle, and the SVD does not exist as a discrete sum. Modern data-driven Koopman SVD methods do not apply here for a fundamental reason.

Stochastic dynamics with a smoothing kernel. When the transition kernel p(xx)p(\mathbf{x}' \mid \mathbf{x}) has a density, conditional expectation averages out fast fluctuations, Kgπgπ\| \mathcal{K} g \|_\pi \le \| g \|_\pi becomes strict on nontrivial observables, and the operator can be Hilbert–Schmidt on L2(π)L^2(\pi). The clean characterization is

KHS2=Eπ[χ2(p(x)π)]+1,\| \mathcal{K} \|_{\mathrm{HS}}^2 = \mathbb{E}_\pi\bigl[ \chi^2\bigl(p(\cdot \mid \mathbf{x}) \,\big\|\, \pi\bigr) \bigr] + 1,

so K\mathcal{K} is Hilbert–Schmidt (in particular compact, hence SVD-admissible) iff the average χ2\chi^2-divergence between the one-step conditional and the stationary distribution is finite (Wu and Noé, 2020; Kostic et al., 2022). Overdamped Langevin dynamics go further: by Weyl’s law on the generator one has σnexp(cn2/dτ)\sigma_n \sim \exp(-c\, n^{2/d}\, \tau), so K\mathcal{K} is trace class and the singular values decay super-polynomially. Stochastic noise regularizes the transition kernel into precisely the regime where SVD-based decompositions are well-defined.

Putting the regimes together: SVD methods for Koopman analysis are stochastic-dynamics methods. Whenever the problem you face is deterministic and measure-preserving, the right object is the unitary spectrum on the circle, not an SVD that does not exist.

When eigenvalues still matter: prediction and pseudospectrum

Closed-form powers. There is one thing the eigenvalue decomposition does that the SVD does not: it diagonalizes powers. If K\mathcal{K} has a complete biorthogonal system Kvi=λivi\mathcal{K} v_i = \lambda_i v_i with duals v~i\tilde v_i, then Ktg=iλitviv~i,g\mathcal{K}^t g = \sum_i \lambda_i^t \, v_i \, \langle \tilde v_i, g \rangle in closed form for every tt, and forecasting at horizon tt becomes multiplication by λit\lambda_i^t on each mode. The SVD has no analogous property under operator composition: (KK)t(\mathcal{K}^* \mathcal{K})^t tells you nothing direct about Kt\mathcal{K}^t unless K\mathcal{K} is normal. This is why DMD and its variants, all centered on eigenestimation, remain the practical forecasting tools even when the operator is highly non-normal.

The ε\varepsilon-pseudospectrum and what eigenvalues miss. For non-normal K\mathcal{K} the eigenvalues alone do not control the dynamics on finite horizons. The relevant object is the ε\varepsilon-pseudospectrum

σε(K)={zC:(KzI)11/ε},\sigma_\varepsilon(\mathcal{K}) = \{ z \in \mathbb{C} : \| (\mathcal{K} - zI)^{-1} \| \ge 1/\varepsilon \},

the set of complex numbers that become eigenvalues under perturbations of size ε\varepsilon. For normal operators it is just an ε\varepsilon-fattening of the spectrum; for strongly non-normal operators it can balloon far into regions with no true eigenvalues at all, and those points govern transient growth and finite-time forecasts (Trefethen and Embree’s Spectra and Pseudospectra, 2005, is the canonical reference). Real dynamical systems routinely violate the textbook “eigenvalues inside the unit disk imply asymptotic decay” picture: continuous spectra coexist with isolated eigenvalues, near-defective (Jordan-block-like) structures inflate finite-time growth far beyond λit|\lambda_i|^t, branch points appear under parameter variation, and modes interact through resonances; in each case the pseudospectrum captures the structure the spectrum alone misses.

ResDMD: residual certification. Computationally, ResDMD (Colbrook, Ayton, and Szőke, 2023) makes this operational: for each candidate eigenpair (λ,v)(\lambda, v) extracted from a data-driven Galerkin approximation, it returns a residual (KλI)v/v\| (\mathcal{K} - \lambda I) v \| / \| v \| that certifies whether the pair is a genuine eigenmode or a discretization artifact, and traces out the pseudospectrum directly when the spectrum is unreliable. The contrast with classical EDMD, whose pseudo-eigenvalues are produced without certification, is the heart of the recent infinite-dimensional spectral-computation program for Koopman analysis.

Parametric Koopman SVD: the computational realization

The practical question is: given trajectory data (xt,xt+1)(\mathbf{x}_t, \mathbf{x}_{t+1}) from a Markov process, how do we compute the top-kk singular subspace of K\mathcal{K} without ever forming the (typically infinite-dimensional) operator and without invoking the numerically unstable SVD or eigendecomposition steps inside a training loop?

This is the question of our NeurIPS 2025 paper, Efficient Parametric SVD of Koopman Operator for Stochastic Dynamical Systems (with Jongha Jon Ryu, Se-Young Yun, Gregory Wornell).

Parametrize the top-kk left and right singular subspaces directly with neural networks f=(f1,,fk)\mathbf{f} = (f_1, \ldots, f_k) and g=(g1,,gk)\mathbf{g} = (g_1, \ldots, g_k). Then minimize the low-rank approximation error

LLoRA(f,g)=Ki=1kfigiHS2\mathcal{L}_{\mathsf{LoRA}}(\mathbf{f}, \mathbf{g}) = \Bigl\| \mathcal{K} - \sum_{i=1}^k f_i \otimes g_i \Bigr\|_{\mathrm{HS}}^2

as the training objective. The LoRA loss has a clean closed form involving only second-moment matrices Eρ0[ff]\mathbb{E}_{\rho_0}[\mathbf{f}\mathbf{f}^\top], Eρ1[gg]\mathbb{E}_{\rho_1}[\mathbf{g}\mathbf{g}^\top], and the joint moment Eρ0(x)p(xx)[f(x)g(x)]\mathbb{E}_{\rho_0(\mathbf{x}) p(\mathbf{x}'\mid\mathbf{x})}[\mathbf{f}(\mathbf{x}) \mathbf{g}(\mathbf{x}')^\top]: no matrix-square-root inverses (which destabilize VAMPnet), no metric distortion regularizer (which limits DPNet), and no numerical SVD inside the training step. The gradient is unbiased under standard minibatch estimation of these moments.

A nesting construction extends this: training the loss j=1kLLoRA(f1:j,g1:j)\sum_{j=1}^k \mathcal{L}_{\mathsf{LoRA}}(\mathbf{f}_{1:j}, \mathbf{g}_{1:j}) produces f\mathbf{f} and g\mathbf{g} whose jjth components span the top-jj singular subspace for each jj, recovering the ordered singular functions rather than just the unordered span. This matters whenever downstream tasks need the dominant mode separately from the second, the second separately from the third, and so on: eigenanalysis of complex molecular dynamics, multi-step prediction with truncation, identification of slow collective variables.

In effect the method takes the operator-theoretic claim (SVD is the right structure for non-normal Koopman; eigendecomposition is the special case) and turns it into a training loop that respects the structure end-to-end. No numerical SVD inside the training step. No assumption of normality. The output is parametric singular functions that scale to high-dimensional stochastic systems where classical DMD variants and their numerically heavy neural extensions stall.

What it amounts to

Most introductions to Koopman methods open with the eigenvalue decomposition. The choice is historical (Koopman 1931 was about Hamiltonian dynamics, where K\mathcal{K} is unitary and the eigenvalue decomposition is the natural object) and pedagogical (eigenvalues are easier to motivate than singular values). For the stochastic, irreversible, non-stationary processes that drive most physical models today, the singular value decomposition is the operator-theoretic primitive that survives non-normality, and methods that train on the right primitive scale where eigen-based methods do not.

The adjoint of the Koopman operator is the backward predictor; reverse-time Koopman is a degenerate case. Normality is the exception, not the rule, and the SVD applies without it.

For the operator-theory background, see the three-part Spectral Theorem series (I, II, III) and the Operator SVD capstone; the Hilbert–Schmidt Koopman operator above sits in the compact regime of Part II.

References
  1. Koopman, B. O. (1931). Hamiltonian systems and transformation in Hilbert space. Proceedings of the National Academy of Sciences 17(5), 315–318.
  2. Eckart, C., and Young, G. (1936). The approximation of one matrix by another of lower rank. Psychometrika 1(3), 211–218.
  3. Trefethen, L. N., and Embree, M. (2005). Spectra and Pseudospectra: The Behavior of Nonnormal Matrices and Operators. Princeton University Press.
  4. Lasota, A., and Mackey, M. C. (2013). Chaos, Fractals, and Noise: Stochastic Aspects of Dynamics. Springer.
  5. Williams, M. O., Kevrekidis, I. G., and Rowley, C. W. (2015). A data-driven approximation of the Koopman operator: extending dynamic mode decomposition. Journal of Nonlinear Science 25(6), 1307–1346.
  6. Mardt, A., Pasquali, L., Wu, H., and Noé, F. (2018). VAMPnets for deep learning of molecular kinetics. Nature Communications 9(1), 5.
  7. Wu, H., and Noé, F. (2020). Variational approach for learning Markov processes from time series data. Journal of Nonlinear Science 30(1), 23–66.
  8. Kostic, V., Novelli, P., Maurer, A., Ciliberto, C., Rosasco, L., and Pontil, M. (2022). Learning dynamical systems via Koopman operator regression in reproducing kernel Hilbert spaces. Advances in Neural Information Processing Systems 35, 4017–4031.
  9. Colbrook, M. J., Ayton, L. J., and Szőke, M. (2023). Residual dynamic mode decomposition: robust and verified Koopmanism. Journal of Fluid Mechanics 955, A21.
  10. Kostic, V. R., Novelli, P., Grazzi, R., Lounici, K., and Pontil, M. (2024). Learning invariant representations of time-homogeneous stochastic dynamical systems. International Conference on Learning Representations.
  11. Jeong, M., Ryu, J. J., Yun, S.-Y., and Wornell, G. W. (2025). Efficient parametric SVD of Koopman operator for stochastic dynamical systems. Advances in Neural Information Processing Systems 38, 25564–25600.

← All posts