15 Wavefronts
Suggestions of all kinds for this book draft are welcome — whether it’s fixing small errors, raising bigger questions, or offering new perspectives. Please share comments through GitHub Issues. To make feedback easier to address, please point to the section you have in mind — by section number or a short snippet of text.
15.1 Wavefronts overview
When we reviewed why Newton’s ray model is clearly wrong, we also realized the need for Huygens’s model of light as a wave Chapter 9. I delayed a full discussion because it wasn’t really necessary to rely on wavefronts until after we finished several other basic principles of optics. Many of the important principles can be described using rays, and three hundred years of practice has taught us that working with rays is useful. Wrong, but useful.
Here we develop more of the infrastructure that is used to describe light as waves. I want you to understand how we use wavefronts in the ISETCam calculations for optics. In the next chapter we will review how we can measure wavefronts using specialized instrumentation (Section 16.2). We will also learn about how adaptive optics systems use wavefront measurements in applications ranging from astronomy (Section 16.3) to ophthalmology (Section 25.4).
15.2 Wavefronts
In Section 10.2 we introduced the connection between the idea of light as a wave and a ray. Recall that given a sketch of a wavefront, we can draw the perpendicular along the wavefront to find the local ray direction Figure 10.2. If a point is relatively far from the aperture, the rays in the incident light field will be nearly parallel (collimated). If the point is close, the rays will be diverging. If the rays have passed through a turbulent medium, they may point in many different directions (Figure 10.2). In each case, the surface perpendicular to these rays is the wavefront.
The wavefront is a description of the light from a single point in the scene. A complete scene is composed of light from many different points, each with its own wavefront. In the following sections, we describe the light from a single point as a wavefront, and then explain why light from different scene points can be analyzed independently.
15.3 Wavefront aberration function
The wavefront aberration is a way to describe how a lens transforms an incident wavefront, just as the point spread function is a way to describe how the lens transforms a point of light in the scene. As the word ‘aberration’ implies, it is the difference between the measured and ideal wavefronts. We measure the difference as the optical path difference between the actual and reference wavefronts at each point in the pupil, typically in units of microns.
Commonly, the reference is a spherical wavefront. Such a wavefront would converge to a sharp image point. Because many pupils are circular—like the pupil in your eye—we parameterize the input and output wavefronts using polar coordinates: \(\rho\) represents the distance from the pupil center and \(\theta\) the angle around the center. We will write the wavefront aberration function as \(W(\rho, \theta)\) 1.
Figure 15.1 illustrates the idea: the gap between the actual wavefront (red) and the ideal wavefront (black) is the wavefront aberration. For a perfect lens, with no aberrations, the actual wavefront coincides with the reference wavefront, and the wavefront aberration is zero: \(W(\rho, \theta)=0\).
15.4 Wavefront description: time course
Figure 15.1 shows the rays as waves rather than straight lines, and indeed they are electric field oscillations traveling away from the point source. A light wave’s electric field in air propagates at the speed of light in that medium, and its wavelength is very short: between 400 and 700 nm for visible light. As a result, the oscillation is at a very high temporal frequency, on the order of \(10^{14}\text{ Hz}\).
\[ E(t) = a(t)\sin(2\pi \nu_0 t + \phi(t)) \tag{15.1}\]
Here \(\nu_0\) is called the carrier frequency, the sine term is the carrier, \(a(t)\) is the wave amplitude, and \(\phi(t)\) is the phase. Natural sources, such as the sun or an incandescent bulb, emit light through countless independent atomic events, so the field’s amplitude and phase fluctuate randomly over time. Hence, the amplitude and phase are not perfectly steady; they fluctuate about the mean randomly, remaining approximately constant over a fairly short interval called the coherence time.
The temporal frequency of a light wave at a fixed position is very high. We can compute it from the relationship between speed, wavelength, and temporal frequency \(\nu\):
\[ c = \nu \lambda \implies \nu = \frac{c}{\lambda} \]
- Visible light wavelength (\(\lambda\)) typically spans roughly \(400\text{ nm}\) to \(700\text{ nm}\) (\(0.4\)–\(0.7\ \mu\text{m}\)).
- The speed of light (\(c\)) in vacuum is approximately \(3 \times 10^8\text{ m/s}\).
For light near the middle of the visible range (\(\lambda \approx 500\text{ nm} = 0.5 \times 10^{-6}\text{ m}\)), this gives:
\[ \nu = \frac{c}{\lambda} \approx \frac{3 \times 10^8\text{ m/s}}{0.5 \times 10^{-6}\text{ m}} \approx 6 \times 10^{14}\text{ cycles/sec (Hz)} \]
The amplitude and phase are random, but they stay correlated only over the coherence time — about a femtosecond (\(10^{-14}\) to \(10^{-15}\text{ s}\)) for typical, broadband visible light (Figure 15.2). On longer time scales, the peaks and troughs of the carrier wash out along with the phase.
15.4.1 Mutual incoherence and independent scene points
Figure 15.2 illustrates two independent light sources centered at \(500\text{ nm}\). Because natural light is emitted through countless independent atomic events, the phase difference between light originating from two different scene points wanders randomly on a timescale of femtoseconds.
Although the instantaneous electric fields from different points superimpose in space (\(E(t) = E_1(t) + E_2(t)\)), their time courses are uncorrelated (orthogonal) over any macroscopic duration. For imaging applications, the integration time of the sensor ranges from tens of microseconds to many milliseconds—trillions of carrier cycles and orders of magnitude longer than the coherence time. Over that integration time, the interference cross-term \(\langle 2 E_1(t) E_2(t) \rangle\) averages strictly to zero.
As a consequence, natural scene points are mutually incoherent: their irradiances add independently on the sensor plane. This fundamental property justifies why we can analyze the optical transformation—the wavefront aberration, the pupil function, and the point spread function—for a single scene point, and then build the full image by summing the irradiances contributed by each point in the scene.
15.4.2 The wavefront representation as a complex number
For a single point source, the light arriving across the pupil is spatially coherent. Because only the time-averaged irradiance (the mean squared amplitude), not the instantaneous electric field oscillation, is measured by the sensor, we can drop the rapid temporal carrier (\(2\pi \nu_0 t\)) and describe the spatial variation of the wave using complex phasor notation. We write the amplitude and phase at each position \(\mathbf{x}\) on the wavefront as
\[ U(\mathbf{x}) = a(\mathbf{x}) e^{i\phi(\mathbf{x})} \tag{15.2}\]
The rate of energy arriving per unit area at a surface is its irradiance (Section 3.1). The irradiance is equal to the squared magnitude of the complex wavefront arriving at that position \(\mathbf{x} = (x, y)\) on the sensor plane:
\[ I(\mathbf{x}) = \vert{}U(\mathbf{x})\vert{}^2 = a(\mathbf{x})^2 \tag{15.3}\]
As you can see, the phase of the wavefront does not influence the amount of energy (irradiance) we record. That depends only on the amplitude2. The irradiance has units of (Watts/m\(^2\)). In imaging and photon-based modeling, irradiance is commonly expressed directly as photon flux density3:
\[ \left[\frac{\text{photons}}{\text{s} \cdot \text{m}^2}\right] \quad \text{or} \quad \left[\frac{\text{q}}{\text{s} \cdot \text{m}^2}\right] \]
15.5 Pupil functions and the PSF
How does a lens transform an incident wave? We asked this same type of question when we considered points of light as input, and this led to the idea of the point spread function (Section 9.2). Here we ask how the lens transforms the input wavefront into an output wavefront. We call the transformation the pupil function, and this idea will be tightly coupled to the idea of the point spread function.
The pupil function is defined by combining the wavefront aberration \(W\) and the amplitude transmission across the exit pupil (or the aperture plane for a simple thin lens):
\[ P(\rho, \theta) = a(\rho, \theta) \, e^{i \, \frac{2\pi}{\lambda} W(\rho, \theta)} \tag{15.4}\]
The pupil function of a distant point, whose rays are collimated, contains the information needed to predict the image. That image is the point spread function.
The complex field in the image plane (at best focus) is proportional to the Fourier transform \(\mathcal{F}\) of the pupil function:
\[ U(x, y) \propto \mathcal{F}\{ P(\rho, \theta) \} \tag{15.5}\]
The PSF is the intensity of this field:
\[ \text{PSF}(x, y) = |U(x, y)|^2 \tag{15.6}\]
The Fourier transform of the pupil function depends on both the aperture opening and the wavefront aberration. Notice that if the optics have no aberrations (\(W = 0\)), the exponential phase term vanishes and the pupil function reduces strictly to the aperture opening \(a\)—a classic result emphasized in Fourier optics courses. But as we will explore in Section 15.8, most imaging systems must account for both aperture diffraction and wavefront phase errors.
15.6 Computations with the pupil function
15.6.1 Paraxial regime and space-invariance
When we introduced the point spread function, we recognized that in some special cases changing the position of the input point only changed the position of the PSF, but not its shape (Section 9.2, Section 38.1).
When we describe lenses using wavefronts, there is a corresponding notion: the paraxial regime. This refers to the range of angles over which planar wavefronts share the same wavefront aberration. Changing the angle of the planar wavefront is like changing the position of the input point; keeping the same wavefront aberration means that the PSF will be unchanged apart from position. In optics and astronomy, the angular region over which the wavefront aberration and PSF remain essentially constant is called the isoplanatic patch. Within this patch, the optical system is space-invariant.
15.7 Modeling wavefront aberrations: Zernike polynomials
The ISETCam software calculations that accompany this book use both wavefront and ray methods4.
The wavefront function \(W(\rho, \theta)\) is often expanded in terms of Zernike polynomials \(Z_n^m(\rho, \theta)\), which form an orthonormal basis over the unit disk (\(0 \le \rho \le 1\), where \(\rho = r / R_{\text{pupil}}\) is the normalized pupil radius):
\[ W(\rho, \theta) = \sum_{n,m} c_n^m Z_n^m(\rho, \theta) \tag{15.7}\]
Here \(n \ge 0\) is the radial order (the degree of the polynomial in \(\rho\)) and \(m\) is the azimuthal frequency (the angular variation, running from \(-n\) to \(+n\) in steps of 2). Each coefficient \(c_n^m\) has units of optical path difference (typically microns, \(\mu\text{m}\)) and weights a specific aberration mode:
- \(n=2, m=0\): defocus
- \(n=2, m=\pm 2\): astigmatism
- \(n=3, m=\pm 1\): coma
- \(n=3, m=\pm 3\): trefoil
- \(n=4, m=0\): primary spherical aberration
This representation enables:
- systematic control over image degradation
- parametric simulations of optical quality
- statistical modeling (e.g., for manufacturing tolerances or biological variation)
When the wavefront aberration is zero (\(W=0\)), the pupil phase is uniform across the aperture. As predicted by the Fourier transform in Equation 15.5, this produces the sharpest possible image for that aperture size: the diffraction-limited Airy pattern.
With defocus aberration, the phase is no longer flat; quadratic curvature appears as concentric phase rings in the pupil. The Fourier transform of this phase distribution yields a broadened, flat-topped PSF with ring structure, drastically reducing peak image contrast.
15.8 The Big Picture
In my classroom experience, students who have taken Fourier transform courses are often struck by the elegant principle that the point spread function (PSF) of a diffraction-limited imaging system can be described directly by the shape of the aperture. It is a neat observation that has the great virtue of being true. In the context of an engineering course, however, this observation can be slippery to integrate because the qualifier “diffraction-limited” slips by so easily. I want to take a few paragraphs to build an intuitive, physical picture integrating what we have learned about apertures, real lenses, wavefronts, and even pinholes. There are crucial subtleties in how these factors interact, and as my wise colleague Stuart Anstis once advised me, “Subtlety doesn’t work.” So let’s not be subtle.
First, the wavefront at the exit pupil and the image intensity are mathematically linked. The complex pupil function (Equation 15.4) describes both the aperture transmission \(a(\rho, \theta)\) and the phase shape of the exiting wavefront. The Fourier transform of this pupil function yields the complex amplitude spread function (Equation 15.5), and its squared magnitude determines the recorded image intensity, which is the PSF (Equation 15.6).
Second, the phase term in the pupil function, \(W(\rho, \theta)\), is defined as the deviation from an ideal spherical wavefront converging precisely onto the sensor plane. Because of this definition, if we change the location of the sensor—moving it away from the conjugate focus (\(d_s \neq d_i\))—the reference sphere changes, and a phase error (defocus) appears in \(W(\rho, \theta)\)5.
Third, when an optical system is truly diffraction-limited, the optics introduce zero wavefront errors, meaning \(W(\rho, \theta) = 0\). In that ideal case, the pupil function reduces strictly to the binary aperture mask, \(a(\rho, \theta)\). This is where the rule in the Fourier transform course originates: in this case the PSF is determined entirely by the Fourier transform of the aperture opening alone.
Fourth, real optical systems rarely remain diffraction-limited across a wide range of conditions. In general, image blur arises from two competing contributions: edge diffraction set by the aperture geometry \(a\), and geometric/optical blur described by phase errors \(W\). The relative importance of the two depends heavily on the operating regime:
- If the aperture is stopped down, phase errors \(W\) are typically tiny across the narrow opening, and edge diffraction \(a\) dominates the blur.
- If the aperture is widened, diffraction blur shrinks, but lens aberrations (or defocus) grow rapidly, causing phase errors \(W\) to dominate the blur.
We encountered a version of this tradeoff earlier when analyzing the simple pinhole camera (Figure 7.4, Section 9.3). For a pinhole of diameter \(D\), an extremely small hole blurs the image through edge diffraction, while an oversized hole blurs the image through geometric projection. The exact same principle governs general lens systems: total blur is an ongoing balance between aperture diffraction and wavefront phase errors.
So enjoy the pure mathematics showing that the aperture shape determines the PSF of a diffraction-limited system. But remember the practical reality: both the aperture boundary and wavefront phase aberrations matter, and we manage them together using the pupil function.
Although it is specified separately for each wavelength of light, \(\lambda\), for now we will not mention wavelength.↩︎
Please excuse me for not explaining why the energy is \(a()^2\) and not \(a()\). It has a good answer that I hope to put into the Resource section some day, involving the electric field of the light and its impact on the electric field within the substrate.↩︎
ISETCam has methods to convert back and forth between these two types of units.↩︎
The ISETCam optics calculations use the wave model. The ISET3d graphics calculations use the ray model.↩︎
Furthermore, the physical distance to the image plane sets the spatial scale of the diffraction pattern on the sensor (\(w \propto \lambda d_i / D\)).↩︎