Spectral Theorem I: adjoint, normality, and the finite-dimensional case
· Revised Jun 01, 2026TL;DR. Building the finite-dimensional spectral theorem from the basis-free adjoint, normality, and Gram–Schmidt induction on eigenvectors. Part I of a three-part series, with a separate Operator SVD capstone.
The spectral theorem is the structural fact that organizes linear algebra on inner-product spaces: every normal operator admits an orthonormal eigenbasis. This post builds the result in finite dimensions: what the adjoint of a linear operator is independently of a chosen basis, what normality really means and what its two parts say about a Markov dynamics, and how Gram–Schmidt induction on the eigenvectors delivers the spectral theorem itself.
Three companion posts continue and apply the result. Spectral Theorem II covers compact operators on a Hilbert space, where eigenvectors still exist but dimension induction does not. Spectral Theorem III covers the general bounded case, where eigenvectors need not exist at all and a projection-valued measure replaces the eigenbasis. The Operator SVD capstone collects the three spectral-theorem results and assembles them, via the polar reduction and the polar form , into the singular value decomposition at each level.
The adjoint and its relation to the transpose
When we write for the transpose of a matrix , we are performing a purely algebraic operation: . The textbook identity , with the dot product on , looks like it identifies the transpose with the adjoint of the operator defined by . It does, but only because both sides of the equation rest on the same implicit pair of conventions: the standard basis and the dot product as the inner product.
Drop either alignment and the two notions diverge. Let with some non-standard basis , and let the inner product on be specified by its Gram matrix . For a linear operator with coordinate matrix in this basis, the geometric adjoint defined by the inner-product identity has coordinate matrix
The transpose still appears, but flanked on both sides by basis-change matrices. The simple identity “adjoint equals transpose” holds only when , i.e., when the chosen basis is orthonormal with respect to the chosen inner product.
The basis-free way to say this is that an inner product on gives a canonical isomorphism
between and its dual . (In the complex case is conjugate-linear; the spirit is the same.) For any linear map between inner-product spaces, the algebraic dual is defined without any inner product by . The geometric adjoint is then the algebraic dual pulled back through the inner-product identifications on the two sides (Halmos, 1958):
This is the object that the matrix formula is computing in coordinates. The construction depends on the inner products on and but not on any choice of basis: change basis, and the matrix of changes by similarity while changes correspondingly; the formula delivers the same operator in the new coordinates.
Why the adjoint is a Hilbert-space notion
In Banach spaces (no inner product), the algebraic dual map still exists, but the geometric adjoint does not. Without an inner-product isomorphism there is no way to land back in . The adjoint is intrinsically a Hilbert-space notion, which is why this post stays in inner-product spaces throughout.
So transpose is an operation on coordinate matrices that depends on a choice of basis. Adjoint is an operation on operators that depends on a choice of inner product but not on a choice of basis. The two coincide whenever the basis is orthonormal in the inner product, which is exactly the default we were using without naming it.
Normality: and what it really says
A linear operator on a finite-dimensional complex inner-product space is normal if . The condition admits several equivalent reformulations, and the one to anchor on says what normality is.
Split into self-adjoint and anti-self-adjoint parts,
both Hermitian by construction. A direct computation gives
where is the commutator. So
The right-hand side is the operator analogue of the trivially commuting decomposition for scalars. For complex numbers, real and imaginary parts commute automatically because scalars commute. For operators, the real and imaginary parts are Hermitian operators that need not commute in general, and normality is the condition that they do. A normal operator is one that “acts like a scalar” in the sense that its self-adjoint and anti-self-adjoint pieces do not interfere with each other.
A second equivalent formulation will do load-bearing work in the spectral-theorem proof below: is normal iff for every . The forward and adjoint orbits of any vector have the same length. As an immediate consequence, if is normal, then and share eigenvectors with conjugate eigenvalues: implies . Apply the norm equality to (which inherits normality from ):
The left side vanishing forces the right side to vanish. This eigenvector-sharing fact is the engine of the spectral theorem, proved two sections below.
Is "normal" basis-independent, and is it the same as a normal matrix?
Two layers of dependence have to be separated. The inner product itself is chosen structure, not read off a norm: every finite-dimensional space admits one (pick a basis, declare it orthonormal), but a given norm induces an inner product only when it obeys the parallelogram law, and topological equivalence of norms supplies no inner product at all. Fix that choice and the rest is clean.
Basis-independent, inner-product-relative. With the inner product fixed, is basis-free, so and are operator equations that hold or fail independently of any basis. They still depend on the inner product: change it and changes, so one operator can be self-adjoint under one inner product and not another.
Not the same as matrix normality, except in an orthonormal basis. In a basis with Gram matrix , the matrix of is . So is self-adjoint iff (equivalently is Hermitian), and normal iff . Only when do these collapse to the familiar and . A Hermitian matrix in a skewed basis need not represent a self-adjoint operator, and the converse fails too.
The invariant content. Matrix normality is preserved by unitary change of basis but not by a general one, so it is not a similarity invariant. General-similarity invariants are the vector-space data (eigenvalues, diagonalizability, Jordan form, no inner product needed); unitary-similarity invariants are the inner-product data (normality, self-adjointness, singular values). Operator normality is exactly the unitary-similarity-invariant content, and is its shadow in orthonormal coordinates. This is why the spectral theorem can state “normal iff an orthonormal eigenbasis exists” without naming a basis.
The reversible and irreversible parts of a Markov dynamics
The split is not just an algebraic device. For the Koopman operator of a Markov process it separates the reversible and irreversible parts of the dynamics, and a small example puts both parts in closed form.
Take a Markov chain on a finite state space with transition matrix , . Its Koopman operator acts on observables by . The inner product that defines the adjoint is the one weighted by the stationary distribution (the row vector with ): . As the opening section stressed, the adjoint depends on this inner product, and a short computation gives
the -weighted transpose. The two Hermitian parts and then have a dynamical reading. is self-adjoint on , so it is the Koopman operator of a reversible chain. measures the failure of to be self-adjoint, and holds exactly when for all , the detailed-balance condition, equivalent to time-reversibility (Norris, 1997). So iff the chain is reversible; when it encodes the net probability current circulating through the state space.
A worked example: the random walk on a 3-cycle. Put three states on a ring. From each state, step clockwise with probability , counterclockwise with probability , and stay with probability . The transition matrix is the circulant
The rotational symmetry forces the stationary distribution to be uniform, , regardless of and : the clockwise/counterclockwise asymmetry changes the current, not the occupancy. Because is uniform, and is the plain transpose, so
Both are circulant, so both are diagonalized by the discrete Fourier basis , , with , and the eigenvalues come out in closed form. The eigenvalues of are
Since and are the Hermitian and anti-Hermitian parts of and share its eigenbasis, their eigenvalues are the real and imaginary parts of the :
Each component has a clean meaning. The mode is the constant observable, fixed by (eigenvalue ) and contributing nothing to . On the two non-trivial Fourier modes, acts with the single real eigenvalue : this is the decay rate, and it is the same for both modes because the reversible part has no preferred sense of rotation. acts with the equal-and-opposite eigenvalues : this is the circulation rate, the irreversible part, and it is precisely the clockwise/counterclockwise asymmetry . A complex eigenvalue of the Koopman operator is therefore a damped oscillation, with the decay supplied by the reversible part and the oscillation supplied by the irreversible part . When the eigenvalues are real, , and the chain is reversible; when the eigenvalues are genuinely complex and the observable rotates as it relaxes.
Because is circulant, and are simultaneously diagonalized by the Fourier basis, so and is normal even when . Normality is therefore strictly weaker than reversibility: the 3-cycle with unequal rates is irreversible (it carries a current) yet still normal. To break normality one has to break the circulant symmetry: a chain with state-dependent hop rates gives , and then no orthonormal eigenbasis exists. The companion post on the Koopman operator takes up the genuinely non-normal case in the stochastic, infinite-state setting.
The spectral theorem via Gram–Schmidt induction
The spectral theorem for normal operators in finite dimensions is now an induction-on-dimension corollary of the eigenvector-sharing fact:
- (Base.) Pick any eigenvalue of (which exists by the fundamental theorem of algebra applied to the characteristic polynomial); let be a unit eigenvector.
- (Step.) The orthogonal complement is -invariant. To see this, use : for , so . The restriction is normal on a space of dimension .
- (Reassembly.) Iterating yields an orthonormal basis of eigenvectors. When eigenvalues are degenerate, the eigenspace itself has higher dimension; one applies Gram–Schmidt orthogonalization within each eigenspace to obtain orthonormal eigenvectors.
Normality versus Schur triangularization
Without normality, Schur’s triangularization theorem still gives a form with unitary and upper-triangular; this is the triangularization available for any operator. Normality is precisely the additional condition under which the off-diagonal entries of vanish, leaving a diagonal matrix. Mechanically, the vanishing happens because of the eigenvector-sharing fact: with , the orthogonal complement of is exactly -invariant rather than approximately so, and the inductive peel-off produces no off-diagonal residue. Without normality the off-diagonal entries record the failure and survive the triangularization (Axler, 2015).
In operator form this is the resolution of identity. Grouping the orthonormal eigenbasis by eigenvalue gives mutually orthogonal eigenspaces that span ; writing for the orthogonal projection onto ,
since acts as on each . The decomposition of the identity into orthogonal projections is what the name resolution of identity refers to; the second equation is the spectral theorem in operator form. In infinite dimensions the sum becomes an integral against a projection-valued measure, which is the subject of Spectral Theorem III.
What it amounts to
The adjoint is the basis-free shadow of the transpose, fixed by the inner product. Normality, , is the condition that the self-adjoint and anti-self-adjoint parts of an operator commute; for the Koopman operator, that the reversible and irreversible parts of a dynamics do not interfere. The spectral theorem turns normality into an orthonormal eigenbasis, written either as a list of eigenvectors or as the resolution of identity above.
All of this is finite-dimensional. The arguments lean on two facts that do not survive the passage to infinite dimensions: the characteristic polynomial has a root, and Gram–Schmidt induction terminates after finitely many steps. Spectral Theorem II rebuilds the result for compact operators on a Hilbert space; eigenvectors still exist, with the variational principle replacing the characteristic polynomial, but the dimension induction has to be replaced by compactness-driven countability and decay. Spectral Theorem III handles the general bounded case, where an operator may have no eigenvectors at all and the functional calculus and projection-valued measure stand in for the eigenbasis entirely. The Operator SVD capstone applies all three results to the polar reduction and reads off the singular value decomposition at each level, with the Eckart–Young best-rank- statement as the standard quantitative consequence.
References
- Halmos, P. R. (1958). Finite-Dimensional Vector Spaces (2nd ed.). Van Nostrand.
- Norris, J. R. (1997). Markov Chains. Cambridge University Press.
- Axler, S. (2015). Linear Algebra Done Right (3rd ed.). Springer.