18 foundation readings sit behind this argument.
The question we were answering:
The idea we opened:
Explanation · Argument synopsis
Back one step Return to the exact step
Reading a graph How does a curve show the relationship between two physical quantities?
The horizontal axis (abscissa) typically shows the independent variable, like position x or elapsed time t. The vertical axis (ordinate) shows the quantity that depends on it, like particle count or probability density p(x).
A flat horizontal line means the quantity does not change across the axis. A rising curve has positive slope (increasing rate), while a symmetric bell-shaped curve shows high concentration in the centre with tails tapering toward zero on both sides.
One worked example Plot particle positions −3, −1, +1, +3 along the horizontal axis with a vertical bar for particle count. The graph shows symmetric spread centered at zero. The average position is 0, but the points are visibly separated from the origin. A stopping point: Check axis labels, units, and whether the curve represents a count, a density, or an accumulated total.
A sign records direction Open this as a full reading page →
Fractions and ratios How does dividing into parts compare two quantities?
A fraction a/b compares quantity a to quantity b, or tells us how many parts of size 1/b make up a. When b is not zero, dividing a by b calculates how many units of b are contained within a.
In physical science, a ratio often carries units: dividing displacement in micrometres by time in seconds produces a rate in micrometres per second. A dimensionless ratio occurs when both quantities share the same units, such as the fraction of an ensemble that moves in a given direction.
One worked example If 1 particle out of 4 moves by 3 units, the proportion is 1/4 = 0.25. If the ensemble doubles to 8 particles and 2 move by 3 units, the proportion is 2/8 = 0.25. Multiplying both the count and the total by 2 preserves the underlying ratio. A stopping point: A ratio compares relative magnitudes; always state whether the physical units cancel or combine.
Adding and averaging Open this as a full reading page →
A sign records direction How can movement add up to zero?
Put the starting point at zero on a ruler. A final mark three units right has displacement +3; one three units left has displacement −3. Adding the signed displacements gives zero. Adding the distances from the start gives six.
Displacement compares the final and initial positions. It is not the length of the path travelled between them.
One worked example Two walkers both finish one unit from their start, one at −1 and the other at +1. The signed average is zero. The average distance is one.
A stopping point: The sign names a direction relative to a chosen axis; it does not mean a negative distance.
Adding and averaging Open this as a full reading page →
Powers of ten and physical units How do we write very large or very small physical measurements compactly?
Microscopic measurements in physics involve scales far outside everyday experience. Writing numbers as a coefficient between 1 and 10 multiplied by a power of ten keeps significant digits clear and avoids counting long strings of zeros.
Units attach physical dimensions to numbers. A distance of 1 micrometre (10⁻⁶ m) squared is 10⁻¹² m², not 10⁻⁶ m². Calculations must keep units alongside numbers through every algebraic step.
One worked example 1 micrometre (1 μm) equals 10⁻⁶ metres (0.000001 m). A typical Brownian displacement of 6 μm after one minute is 6 × 10⁻⁶ m. Squaring this displacement gives (6 × 10⁻⁶ m)² = 36 × 10⁻¹² m² = 3.6 × 10⁻¹¹ m². A stopping point: Every physical number must carry its units, and powers of ten apply to both coefficient and unit.
Fractions and ratios Open this as a full reading page →
Squares and square roots Why does four times a mean square mean twice the RMS?
Both +3 and −3 have square 9. Squaring therefore removes the sign while giving greater weight to a larger magnitude. A square root returns the result to the original kind of unit: the square root of a square micrometre is a micrometre.
The square root of four q is twice the square root of q, for nonnegative q.
One worked example The mean square of −3, −1, +1, +3 is 5. Doubling each displacement gives −6, −2, +2, +6, whose mean square is 20. The RMS changes from approximately 2.236 to 4.472: twice as large, not four times as large.
A stopping point: Use the nonnegative square root when reporting a magnitude.
Adding and averaging Open this as a full reading page →
Adding and averaging What does an average keep, and what does it lose?
Imagine four equally sized containers holding 3, 1, 1 and 3 units. Combining them gives 8 units. Sharing the total equally between the four containers gives 2 in each. Two is the average; the original containers were not all equal.
The number of observations and the sum of their values are different quantities. Doubling the number of repeated observations doubles the sum but leaves the average unchanged.
One worked example Add 3 + 1 + 1 + 3 to obtain 8. Count four values. Divide 8 by 4 to obtain 2. Repeating the list gives 16 divided by 8, still 2. A stopping point: An average is a total shared equally among the number of observations.
Open this as a full reading page →
Rates of change and derivatives What physical rate does a derivative measure?
When a quantity changes over an interval of time or space, the ratio of the change in output to the change in input gives an average rate. As the measurement interval shrinks toward zero, this average ratio approaches a definite limit: the instantaneous derivative.
Geometrically, the derivative is the slope of the tangent line to the function graph at a single point. Physically, it always carries the units of the output divided by the units of the input. For position over time, the derivative is velocity in metres per second; for concentration over position, it is a spatial gradient in particles per metre to the fourth power.
In 1905, Einstein used derivatives to relate microscopic flux to macroscopic concentration gradients, and to extract temperature and wavelength dependencies from radiation laws without guessing unmeasured intermediates.
One worked example The derivative of f with respect to x is the limit of the difference quotient as delta x approaches zero.
If position x(t) = c t^2 with c = 3 metres per second squared, the change between t and t + Δt is c(t+Δt)^2 - ct^2 = 2ctΔt + c(Δt)^2. Dividing by Δt gives 2ct + cΔt. In the limit Δt -> 0, the instantaneous velocity is exactly 2ct = 6t metres per second.
A stopping point: A derivative is an instantaneous rate bearing explicit units; it is not a fraction of two separate isolated zeros.
Interactive construction: local sensitivity and derivative units In section 8 of the light-quanta paper, Einstein predicts that when light liberates electrons from a cathode, increasing the light frequency ν increases the required stopping potential V linearly. The derivative dV/dν is the local sensitivity of stopping voltage to incident frequency.
Choose a frequency step Δν:
+1.0 × 10¹⁴ Hz (large step) +5.0 × 10¹³ Hz (medium step) +1.0 × 10¹³ Hz (small nudge) +2.0 × 10¹² Hz (fine nudge)
Observed sensitivity response Baseline frequency (ν₀): 6.00e+14 Hz Frequency nudge (Δν): +1.00e+13 Hz Potential change (ΔV): +4.135668e-2 V Sensitivity ratio (ΔV / Δν): 4.135667696e-15 V·s (or V/Hz) Universal ratio h/e: 4.135667696e-15 V·s Summary of frequency steps and sensitivity ratio Nudge size Δν (Hz) ΔV (V) Ratio ΔV / Δν (V·s) +1.0 × 10¹⁴ Hz 1.0e+14 4.1357e-1 4.135668e-15 +5.0 × 10¹³ Hz 5.0e+13 2.0678e-1 4.135668e-15 +1.0 × 10¹³ Hz 1.0e+13 4.1357e-2 4.135668e-15 +2.0 × 10¹² Hz 2.0e+12 8.2713e-3 4.135668e-15
Textual summary of the construction A derivative is not a dimensionless number; it has physical units determined by the ratio of output units to input units. Here, dividing volts by hertz yields volt-seconds. Regardless of how small the nudge step Δν is chosen, the ratio ΔV / Δν evaluates to the exact physical constant h/e ≈ 4.14 × 10⁻¹⁵ V·s, confirming that the sensitivity of stopping potential to frequency is universal and independent of the metal.
Functions and graphs Open this as a full reading page →
Density is not probability What does the height of a probability curve mean?
Divide a ruler into bins. The fraction of observations in a bin estimates its probability. Dividing that fraction by the bin width estimates a density. Changing the ruler from metres to micrometres changes the numerical density but not the probability of the same physical interval.
The probability of the interval from a to b is the integral of the density over that interval.
A continuous distribution gives probability zero to any single exact position. At the starting time the ideal point source is instead an atom of probability one, not a finite curve with infinite height.
One worked example A uniform density of 0.25 per micrometre on a four-micrometre interval has total probability one. A one-micrometre subinterval has probability 0.25; a two-micrometre subinterval has probability 0.5.
A stopping point: Always distinguish curve height, interval area and an individual observation.
Adding and averaging Open this as a full reading page →
Forces on charges, currents, and electromagnetic waves How do charges interact with electric and magnetic fields?
Electric charge is the intrinsic property of matter that produces and responds to electromagnetic fields. A charge q at rest in an electric field E experiences an electrostatic force F = q E. When moving with velocity v through a magnetic field B, the charge experiences an additional perpendicular magnetic force F = q (v x B). Together, these form the Lorentz force.
Moving a charge through an electric potential difference Delta V transfers potential energy Delta W = q Delta V. For a fundamental electron charge e = 1.602e-19 C accelerated across a potential of 1 V, the energy gained is defined as 1 electron-volt (1 eV = 1.602e-19 J).
In microscopic matter, an electron bound to an equilibrium position by a restoring force acts as a harmonic oscillator or resonator. When disturbed by incoming electromagnetic waves, it absorbs and reradiates energy at its natural resonant frequency nu_0. Einstein's 1905 light-quanta paper opened with Planck's model of such resonators in thermal equilibrium with radiant energy.
Maxwell's electrodynamics expresses four physical laws in words: electric charges act as sources of electric flux; magnetic field lines are closed loops with no isolated magnetic charges; a time-varying magnetic field induces a circulating electric field (Faraday induction); and electric currents alongside time-varying electric fields generate circulating magnetic fields (Maxwell-Ampere law). Combined, these equations govern self-propagating electromagnetic waves traveling at speed c in vacuum.
Einstein pointed out in 1905 that classical electrodynamics treated the relative motion of a magnet and a conductor with an artificial asymmetry: moving the magnet created an electric field in space that drove current, whereas moving the conductor created no electric field but rather a magnetic Lorentz force on electrons. Yet the physical current was identical in both descriptions.
One worked example Force equals charge times the sum of electric field and the cross product of velocity with magnetic field, and work equals charge times potential difference.
Accelerating an electron of charge e = 1.602e-19 Coulombs across an electric potential difference of Delta V = 1.0 Volt gives kinetic energy Delta W = (1.602e-19 C)(1.0 V) = 1.602e-19 Joules = 1 eV. An electron bound with effective spring constant k_s and mass m oscillates at natural frequency nu_0 = (1 / 2 pi) sqrt(k_s / m).
A stopping point: Charge is the property that makes electric forces; fields mediate force between separated charges without action-at-a-distance.
Functions and graphs Rates of change and derivatives Open this as a full reading page →
Entropy, temperature, and a stated constraint Why does inverse temperature describe entropy gained per added energy?
Entropy is a state quantity. For a reversible transfer of heat into a system at absolute temperature T, the transferred entropy is heat divided by T. At fixed volume, with no other work or exchanged matter, that heat transfer changes the internal energy.
At fixed volume and the stated constraints, entropy change is energy change divided by absolute temperature.
The subscript matters: allowing work from changing volume introduces an additional term. An entropy derivative describes a local change along specified constraints, not every possible process.
When energy is shared among equilibrium subsystems, a redistribution that conserves total energy cannot increase the already maximized entropy. The entropy slopes are therefore equal, so the temperatures agree.
Open the foundation: Partial derivatives and held-fixed quantities
One worked example An ideal reservoir remains at 300 kelvin while it receives 3 joules reversibly. Its entropy increases by 0.01 joule per kelvin. For a finite body whose temperature changes, integrate dE/T(E) instead of dividing by an arbitrarily selected temperature.
Two proposed entropy functions differing by a constant have the same derivative. A boundary condition is needed to select between them. If the unresolved quantity is an entropy density, multiplying it by different volumes can make that constant matter to a total entropy difference.
A stopping point: State what is held fixed, distinguish total entropy from entropy density, and supply an integration condition before interpreting a volume-dependent difference.
Partial derivatives and held-fixed quantities Adding continuously Open this as a full reading page →
Exponentials and continuous scaling Why do exponential functions govern thermal distributions and random diffusion?
The exponential function e^u, often written exp(u), is the unique mathematical function whose derivative with respect to its argument equals the function value itself. Whenever a physical growth or decay rate is proportional to the amount already present, the resulting trajectory is an exponential.
A fundamental rule of dimensional physics is that the argument of any transcendental function, including the exponential, must be a dimensionless number. In Wien's radiation law exp(-beta nu / T), the product beta nu has dimensions of temperature, so dividing by temperature T yields a pure dimensionless ratio. In the Brownian spreading Gaussian exp(-x^2 / (4Dt)), x^2 has units of squared metres and 4Dt has units of (m^2/s)*s = m^2, ensuring the exponent is dimensionless.
Exponentials also arise naturally from the multiplication of independent probabilities across repeated random steps or subdivided spatial cells.
One worked example Density f of x and t equals one over square root of four pi D t times exponential of minus x squared over four D t.
At the center x = 0, the exponential factor exp(0) = 1. At one standard deviation x = sqrt(2Dt), the exponent is -1/2 and exp(-0.5) approx 0.6065. At x = sqrt(4Dt), the exponent is -1 and exp(-1) approx 0.367879, showing symmetric bell-shaped decay.
A stopping point: The argument inside an exponential function must always be a dimensionless pure number.
Interactive construction: repeated proportional changes and dimensionless exponents In linear change, an equal amount is added in every equal interval: y = y₀ + mt. In exponential change, the quantity is multiplied by an equal factor in every interval: y = y₀ · rⁿ. Because each change is proportional to the current amount, exponential functions naturally describe continuous growth and decay.
Select a step index n to inspect compounding:
Step n = 0 (t = 0s) Step n = 1 (t = 10s) Step n = 2 (t = 20s) Step n = 3 (t = 30s) Step n = 4 (t = 40s) Step n = 5 (t = 50s)
Inspection at step n = 1 Formula: f(1) = (e−0.5 )1 = e−0.5
Fraction remaining: 0.606531 (60.65%)
Each 10-second interval scales the previous value by exactly e−0.5 ≈ 0.606531.
Compounding decay steps with constant multiplier e⁻⁰·⁵ Step n Elapsed time t (s) Step multiplier Fraction remaining Percent 0 0 1.000000 1.000000 100.00% 1 10 e⁻⁰·⁵ ≈ 0.606531 0.606531 60.65% 2 20 e⁻⁰·⁵ ≈ 0.606531 0.367879 36.79% 3 30 e⁻⁰·⁵ ≈ 0.606531 0.223130 22.31% 4 40 e⁻⁰·⁵ ≈ 0.606531 0.135335 13.53% 5 50 e⁻⁰·⁵ ≈ 0.606531 0.082085 8.21%
Why exponents must always be dimensionless You cannot evaluate e raised to three meters or five seconds, because the series definition eu = 1 + u + u²/2! + … would require adding meters to square meters. In every physical law, dimensional quantities inside exponents are strictly cancelled by matching units:
Wien’s law (Light §4): e−βν/T . The frequency ν (s⁻¹) and temperature T (K) are balanced by β = h/k (s·K), making βν/T dimensionless.Brownian diffusion Gaussian (Brownian §4): e−x²/(4Dt) . The numerator x² has units m²; the denominator 4Dt has units (m²/s · s) = m², so x²/(4Dt) is purely dimensionless.Textual summary of the construction Equal steps in the independent variable produce equal multiplicative ratios in the dependent variable. The characteristic scale τ sets the interval over which the quantity changes by a factor of 1/e ≈ 0.367879. All physical exponents are dimensionless ratios of the independent variable to this characteristic scale.
Rates of change and derivatives Logarithms and product-to-sum relations Open this as a full reading page →
Fields, continuous waves, and harmonic functions What is a continuous wave field and how is its energy measured?
A field assigns a definite physical quantity to every location in space and time. A scalar field assigns a single number, such as temperature or particle density, while a vector field assigns a magnitude and direction, such as electric field strength E or magnetic field B.
Harmonic waves propagate oscillating field values through space according to the phase factor kx - omega t, where the wavenumber k = 2 pi / lambda relates to wavelength lambda and angular frequency omega = 2 pi nu relates to cyclic frequency nu. The wave crests travel at the phase speed c = lambda nu.
Because optical frequencies oscillate hundreds of trillions of times per second (for example, 600 THz corresponds to green-cyan light with wavelength lambda = c / nu = 499.65 nm), measuring instruments record the time-averaged energy flux rather than the instantaneous field oscillation. Over any complete period, the average of cos^2(omega t) is exactly 0.5.
When an isotropic source emits total power P into three dimensions, the energy spreads evenly over concentric spherical wavefronts of area 4 pi r^2. The intensity at distance r is given by I = P / (4 pi r^2). For a 1 W point source, the intensity is 0.0795775 W/m^2 at distance r = 1 m, and drops to one-quarter (0.0198944 W/m^2) at distance r = 2 m.
The Doppler shift in sound is an illustrative analogy for wave frequency shifts: sound waves travel through a material medium (air), which creates an asymmetry between a moving source and a moving listener. This analogy has a strict physical limit: light requires no material ether medium, and in 1905 Einstein demonstrated that electromagnetic wave transformation depends purely on relative velocity between observers.
One worked example Intensity equals total power divided by four pi r squared, and the time average of cosine squared over a period is one half.
At frequency nu = 600 THz (6.0e14 Hz) with light speed c = 299792458 m/s, the wavelength is lambda = c / nu = 499.65 nm. A 1 W source produces intensity I = 1 / (4 pi (1)^2) = 0.0795775 W/m^2 at 1 m, and I = 1 / (4 pi (2)^2) = 0.0198944 W/m^2 at 2 m.
A stopping point: A field is a value assigned to every place; observed optical intensity is a time average, not an instantaneous pulse.
Functions and graphs Rates of change and derivatives Open this as a full reading page →
Functions and graphs How does a mathematical function record relationships between physical quantities?
A physical function is an unambiguous rule that assigns to each value of an independent input (such as elapsed time t or coordinate position x) a specific value of a dependent quantity (such as particle displacement, local concentration, or field energy). It describes how one aspect of a physical system responds when another changes.
Plotting a function creates a curve on a coordinate plane. The horizontal axis represents the input, and the vertical axis represents the output. Every point on the curve represents a simultaneously paired measurement. Because both axes correspond to physical measurements, any slope or area derived from the graph carries the combined dimensions of those axes.
Distinguishing functional laws from empirical scatter is essential. A theoretical function predicts the ideal expectation of an ensemble, such as the mean square displacement growing linearly with time. Individual experimental trials fluctuate around this expectation, but the functional form governs the ensemble average.
One worked example Mean square coordinate displacement equals two times the diffusion coefficient times time.
For a diffusion coefficient D = 0.5 square micrometres per second, plotting mean square displacement against time t produces a straight line through the origin with slope 2D = 1.0 square micrometres per second. At t = 1 second the value is 1.0 square micrometre; at t = 4 seconds it is 4.0 square micrometres.
A stopping point: A graph is a record of paired quantities with physical units; a slope without units is not an explanation.
Interactive construction: from measurement table to curve In section 5 of the Brownian motion paper, Einstein shows that while the mean displacement is zero, the mean squared displacement grows linearly with time: ⟨x²⟩ = 2Dt. Consequently, the observable root-mean-square displacement grows as the square root of time: √⟨x²⟩ ∝ √t.
Plot next point (5/5) Plot all Reset Showing 5 of 5 points plotted. (0s, 0µm) (1s, 2µm) (4s, 4µm) (9s, 6µm) (16s, 8µm) Time t (seconds) RMS displacement (µm) Dashed line: continuous √t trajectory. Dots: discrete observations. Table of paired measurements (Brownian §5) Time t (s) ⟨x²⟩ (µm²) √⟨x²⟩ (µm) Status 0 0 0.0 Plotted 1 4 2.0 Plotted 4 16 4.0 Plotted 9 36 6.0 Plotted 16 64 8.0 Plotted
Textual summary of the construction Each observation pairs an elapsed time in seconds with an accumulated squared displacement in square micrometers. Plotting these pairs demonstrates that displacement does not scale proportionally with time (which would indicate constant drift velocity), but rather with the square root of time (the hallmark of diffusive random walks).
Reading a graph Open this as a full reading page →
Adding continuously What operation does an integral describe here?
Approximate the area under a curve by rectangles. Each contributes height times width. Add all contributions, then refine the widths. When those sums approach a limit, that limit is the integral.
For a probability density, a rectangle has units of inverse length times length, leaving a dimensionless probability. For total probability the whole area must be one.
Integration by parts comes from adding the product rule for differentiation over an interval: the integral of u times the derivative of v equals the endpoint product minus the integral of v times the derivative of u.
One worked example Integration by parts moves a derivative from one factor to another and retains the endpoint term.
A constant density of 0.25 per micrometre over 4 micrometres gives 0.25 times 4 = 1, regardless of how many equal rectangles we use.
A stopping point: The width and the endpoint term are part of the calculation, not decorations on the integral sign.
Density is not probability Open this as a full reading page →
Logarithms and product-to-sum relations Why do logarithms transform multiplicative probabilities into additive thermodynamic quantities?
The natural logarithm ln(x) is the inverse of the exponential function: ln(e^u) = u and e^(ln x) = x. Its defining algebraic property is that it converts products into sums: ln(A * B) = ln(A) + ln(B), and powers into products: ln(f^n) = n * ln(f).
In 1905 (§5 of the light-quanta paper), Einstein proved that if entropy S is an additive state function for independent systems (S = S_1 + S_2) while the statistical state probability W is multiplicative (W = W_1 * W_2), then the connection S = phi(W) must satisfy phi(W_1 * W_2) = phi(W_1) + phi(W_2). The only continuous solution is the logarithm: S = k * ln(W) + constant.
For n = 10 independent molecules each occupying a volume fraction f = 1/2, the joint probability is (1/2)^10 and its logarithm is 10 * ln(1/2) approx -6.931472. Inverting Wien's radiation law u = A nu^3 exp(-beta nu / T) to solve for inverse temperature requires taking the logarithm: ln(u / (A nu^3)) = -beta nu / T.
Notation note: In 1905 German scientific literature (including Annalen der Physik), 'lg' denoted the natural logarithm with base e. Modern ISO notation reserves 'lg' for the common base-10 logarithm log_10 and uses 'ln' for the natural logarithm. For example, 1905 printed 'lg 2' meant ln(2) approx 0.693147, not log_10(2) approx 0.301030.
One worked example Entropy of combined independent systems equals Boltzmann constant times natural log of product of state weights, which equals sum of individual entropies.
Natural log of 2 is ln(2) approx 0.693147, whereas common log of 2 is log_10(2) approx 0.301030. For W_1 = 4 and W_2 = 8, W = 32: ln(4) approx 1.386294, ln(8) approx 2.079442, and ln(32) approx 3.465736 = 1.386294 + 2.079442.
A stopping point: The natural logarithm is the unique continuous function mapping independent product states to additive thermodynamic quantities; 1905 printed 'lg' denotes the natural logarithm.
Interactive construction: turning multiplication into addition In section 5 of the light-quanta paper, Einstein reasons about the entropy S of independent systems. When two independent systems with microstate counts W₁ and W₂ are combined, the total number of configurations multiplies: W = W₁ · W₂. However, the thermodynamic entropy must add: S = S₁ + S₂. The only continuous function satisfying φ(W₁ · W₂) = φ(W₁) + φ(W₂) is the logarithm: S = k ln W + const.
Notation display toggle:
Notation: Modern ISO standard (ln = natural log) Click or press Enter to toggle between 1905 historical print and modern ISO symbols.
1905 historical notation vs Modern ISO standard In 1905 German scientific printing (including Annalen der Physik ), the symbol lg denoted the natural logarithm (base e).
In modern ISO 80000-2 notation, ln denotes the natural logarithm, while lg is reserved for the common base-10 logarithm (log₁₀).
1905 printed “lg 2”: 0.693147 (natural logarithm ln 2) Modern ISO “lg 2” (log₁₀ 2): 0.301030 (common base-10 logarithm)
Verification of logarithmic product-to-sum identity: ln(W₁ · W₂) = ln(W₁) + ln(W₂) State W₁ State W₂ Product W₁ · W₂ ln(W₁) ln(W₂) Sum ln(W₁) + ln(W₂) ln(W₁ · W₂) 2 4 8 0.693147 1.386294 2.079442 2.079442 3 5 15 1.098612 1.609438 2.708050 2.708050 10 10 100 2.302585 2.302585 4.605170 4.605170
Textual summary of the construction The table demonstrates that for any pair of numbers, the logarithm of their product exactly equals the sum of their individual logarithms. This algebraic homomorphism bridges statistical mechanics (where independent configurations multiply) and macroscopic thermodynamics (where entropy is an extensive, additive quantity). When reading 1905 papers, readers must translate printed “lg” to natural “ln” to obtain the correct physical entropies.
Exponentials and continuous scaling Open this as a full reading page →
Partial derivatives and held-fixed quantities What does a partial derivative mean when several variables change simultaneously?
Many physical quantities depend on more than one parameter: particle density depends on both position x and time t; gas entropy depends on both volume V and temperature T. When asking how such a quantity changes, one must specify which variable is moving and which variables are being held constant.
The partial derivative ∂f/∂x represents the rate of change of f with respect to x while holding time t strictly fixed. Conversely, ∂f/∂t represents the rate of accumulation at a fixed position x over time. In thermodynamics, holding temperature fixed produces an isothermal derivative, while holding entropy or volume fixed produces an adiabatic or isochoric derivative.
Einstein's derivation of the diffusion equation equates the time rate of accumulation at a fixed location to the divergence of spatial flux: ∂f/∂t = D ∂^2f/∂x^2. Both sides describe rates under different held-fixed constraints.
One worked example Spatial partial derivative of density with respect to x holding time t fixed.
Example 1 (Time fixed when moving through space): For diffusion profile f(x, t) = (4 pi D t)^(-1/2) exp(-x^2 / (4Dt)), the spatial gradient quantifies concentration variation along a channel. Here, elapsed time t is strictly fixed as the held-fixed parameter. A spatial snapshot taken at one fixed instant evaluates to ∂f/∂x = -x/(2Dt) f(x, t).
Time partial derivative of density with respect to time holding position x fixed.
Example 2 (Position fixed when tracking time): The time accumulation rate records concentration changes at a fixed location. Here, spatial position x is strictly fixed as the held-fixed parameter. A stationary probe at one coordinate records accumulation rate ∂f/∂t = (-1/(2t) + x^2/(4Dt^2)) f(x, t). Equating accumulation to spatial flux divergence yields ∂f/∂t = D ∂^2f/∂x^2.
Isothermal volume derivative with respect to pressure holding temperature T fixed versus adiabatic volume derivative holding entropy S fixed.
Example 3 (Thermodynamic derivatives: isothermal versus adiabatic): In gas thermodynamics, the volume response to pressure depends on thermal boundary conditions. The isothermal derivative (∂V/∂p)_T explicitly names temperature T as the held-fixed quantity while heat exchanges freely with a bath. Conversely, the adiabatic derivative (∂V/∂p)_S explicitly names entropy S as the held-fixed quantity under thermal insulation. Because isothermal compression permits heat release, the isothermal compressibility exceeds the adiabatic compressibility. Writing ∂V/∂p without identifying the held-fixed parameter is physically incomplete.
Monochromatic radiation entropy density derivative with respect to energy density holding volume V and frequency nu fixed.
Example 4 (Radiation entropy derivative in Light Quanta §3): For blackbody radiation at frequency nu, the entropy density s_nu depends on both spectral energy density u_nu and enclosure volume V. Einstein's derivation of Wien's displacement law takes the partial derivative (∂s_nu/∂u_nu)_(V, nu) = 1/T. Here both cavity volume V and radiation frequency nu are strictly held fixed as the held-fixed parameters while varying energy density u_nu.
Chain rule transformation of spatial partial derivative holding rest-frame coordinates fixed into moving-frame partial derivatives.
Example 5 (Transformed derivatives in Relativity §6): In coordinate transformations between a resting frame (x, y, z, t) and a moving frame (x', y', z', t'), the partial derivative (∂/∂x)_(y, z, t) explicitly holds resting coordinates y, z, and time t fixed. When expressed in moving coordinates via the relativistic chain rule, it becomes a combination of moving spatial derivative (∂/∂x')_(y', z', t') holding y', z', and t' fixed, and moving time derivative (∂/∂t')_(x', y', z') holding x', y', and z' fixed. Changing coordinate systems changes which physical quantities are held fixed.
A stopping point: A partial derivative is mathematically and physically undefined until every held-fixed parameter is explicitly identified.
Interactive construction: what is held fixed in a partial derivative In sections 3 and 4 of the Brownian motion paper, Einstein tracks the concentration of suspended particles c(x, t) as a function of both position x along a tube and elapsed time t. Because two independent variables can change, asking for “the rate of change of concentration” is ambiguous until you specify which coordinate is held fixed.
Choose which coordinate to hold fixed:
Hold time t fixed: Spatial derivative ∂c/∂x Hold position x fixed: Time derivative ∂c/∂t
Case A: Hold time t fixed (∂c / ∂x) Quantity held fixed: Time t (a single snapshot across the tube).
Meaning: Spatial concentration gradient. You inspect different positions along the tube at one frozen instant.
Physical units: particles / (µm³ · µm) = particles / µm⁴.
Role in Brownian motion: Fick’s first law of diffusion states that the particle flux is proportional to this spatial gradient: J = −D (∂c/∂x).
Thermodynamic examples: how the fixed constraint changes the derivative In thermodynamics, the same symbols have completely different numerical values depending on what is held fixed:
Thermodynamic partial derivatives and their held-fixed constraints Process Derivative Quantity held fixed Physical behavior Isothermal (∂p / ∂V)T Temperature T fixed Heat flows in or out to maintain constant temperature Adiabatic (∂p / ∂V)S Entropy S fixed (no heat exchange) Gas warms upon compression; stiffer response than isothermal Isochoric (∂p / ∂T)V Volume V fixed Rigid vessel; pressure rises directly with heating
Textual summary of the construction Writing ∂c/∂x asserts that t is held constant during differentiation. Writing ∂c/∂t asserts that x is held constant during differentiation. These two operations describe different physical phenomena and carry different physical dimensions. In thermodynamics, the subscript notation (∂p/∂V)T versus (∂p/∂V)S makes this essential distinction visible on the page.
Rates of change and derivatives Open this as a full reading page →
Probability and independence Why do the cross terms disappear?
For independent choices, the probability of one result together with another is the product of their probabilities. Independence says learning the first result does not change the probabilities of the second. It does not mean that the second result must undo the first.
The expected product of independent centred variables A and B is zero.
For nonzero means the product of the means survives. Correlations can also preserve a cross term. The diffusion argument must state which case it assumes.
One worked example For two fair independent steps of size one, list (++), (+−), (−+), (−−). Each has probability one quarter. Their products are +1, −1, −1, +1. The average product is zero. Their sums are +2, 0, 0, −2. The squared sums are 4, 0, 0, 4, with average 2. If the second step always repeats the first, the squared sum is always 4. That is a different, correlated model. A stopping point: Ask whether learning one result changes the probabilities for the next.
Adding and averaging A sign records direction Open this as a full reading page →
Energy of motion and inertia What can a change in energy of motion tell us when speed stays the same?
Work transfers energy when a force acts through a displacement. The energy associated with the motion of a body is called kinetic energy. It is not the same as the energy of its internal heating, chemistry, or other stored processes.
For the same body at ordinary slow speeds, doubling speed multiplies its energy of motion by four. At a fixed speed, doubling inertial mass doubles that energy. The word inertia describes resistance to a change of motion, not a measurement of gravitational weight.
In Newtonian mechanics, kinetic energy is one half times inertial mass times speed squared.
The formula is the low-speed rule. It cannot be assumed exact for a traveler moving at a substantial fraction of light speed. To identify inertia in a relativistic comparison, examine the coefficient as speed approaches zero.
One worked example Compare two bodies at 2 metres per second. In the Newtonian model a body of 3 kilograms has 6 joules of energy of motion; a body of 2 kilograms has 4 joules. The speed is unchanged, but the energy of motion differs by 2 joules.
At the same low speed, the kinetic-energy drop equals one half times the mass decrease times the squared speed.
This example assumes independently specified masses; it is not evidence for mass–energy equivalence. The mass-energy argument instead calculates an energy difference from emitted light and then identifies its low-speed coefficient.
A stopping point: Less energy of motion at the same low speed means a smaller inertial mass within the Newtonian approximation. The comparison alone does not supply an absolute internal energy.
Squares and square roots Fractions and ratios Open this as a full reading page →