Skip to the chapter

Annus Mirabilis · saved explanation preview

Brownian motion: from wandering to a measurable law: §2 · Osmotic pressure from the molecular-kinetic theory

Über die von der molekularkinetischen Theorie der Wärme geforderte Bewegung von in ruhenden Flüssigkeiten suspendierten Teilchen

This is newly written explanation in modern notation, and its editorial review is pending. It is not the German source, an English translation, or a complete edition of the paper. The German source face holds a machine-drafted transcription with hand correction that no one has reviewed yet, and the facsimile face shows the pinned journal pages; an English translation is not ready. The headings name the part of the argument each passage discusses; they are not a list of the paper’s paragraphs.

This file contains the available authored explanation, not a reviewed transcription or translation. No simulation, search index, notebook, or private settings are included. Only explicit online links need a connection.

Current online chapter

Choose a reading

§2 · Osmotic pressure from the molecular-kinetic theory

How molecular theory gives the osmotic law without solving the motion

How can a theory of countless moving molecules give the osmotic pressure of dissolved molecules and suspended bodies without anyone solving their motion?

A footnote to the heading says what this section assumes and what it is for. It takes as known Einstein's papers on the foundations of thermodynamics (Ann. d. Phys. 9, p. 417, 1902; 11, p. 170, 1903), and says that neither those papers nor this section is needed to understand the results of the present paper. The section is still worth reading, because it answers the question §1 leaves open: how can a theory of countless colliding molecules give a law as simple as van 't Hoff's, for dissolved molecules and suspended bodies alike, without anyone solving their motions?

Einstein describes the whole system, liquid, wall and particles, by state variables p1,…,plp_1, \ldots, p_l that fix its momentary state completely, for example the coordinates and velocity components of all its atoms. How they change in time is given by equations of the form

∂pν∂t=φν(p1,…,pl)\frac{\partial p_\nu}{\partial t}=\varphi_\nu(p_1,\ldots,p_l)

The rate of change of p nu with time equals phi nu, a function of p 1 through p l.

with the condition ∑∂φν∂pν=0\sum \frac{\partial \varphi_\nu}{\partial p_\nu} = 0. Two of these letters are used elsewhere in the paper for other things: these pνp_\nu are not the osmotic pressure p of §1, and these φν\varphi_\nu, the rates of change of the state variables, are not the φ of §4.

For such a system his earlier theory gives the entropy S as an expression containing the logarithm, printed lg and meaning the natural logarithm, of an integral taken over every combination of the state variables that the conditions of the problem allow. In it T is the absolute temperature, Ē (printed with a bar) the energy of the system, and E the energy as a function of the pνp_\nu. The constant is printed as 2κ, and Einstein ties κ to N by 2κN = R, so 2κ is R/N. For the free energy F he obtains

F=−RNTlg⁡∫e−ENRT dp1…dpl=−RTNlg⁡BF=-\frac{R}{N}T\lg\int e^{-\frac{EN}{RT}}\,dp_1\ldots dp_l=-\frac{RT}{N}\lg B

F equals minus R over N times T times the logarithm of the integral of e to the minus E N over R T, over d p 1 through d p l, which equals minus R T over N times the logarithm of B.

and the integral is named B.

Return to the free energy as a logarithm of the integral B. (included in this file)

Even if the molecular picture were fixed in every detail, Einstein says, computing B would be so hard that an exact calculation of F is hardly conceivable. But the pressure needs only how F depends on the volume V∗V^* in which all the particles are held. (Particles, 'Teilchen', is his short word for dissolved molecules and suspended bodies alike.)

Put n particles in V∗V^*, held there by a semipermeable wall, their total volume small compared with V∗V^*. Where the wall stands limits the range of the integral B. Name the coordinates of the particles' centres of gravity x1,y1,z1x_1, y_1, z_1 through xn,yn,znx_n, y_n, z_n, give each centre a tiny box inside V∗V^*, and ask for the part of B that comes from states with every centre in its box. It has the form

dB=dx1 dy1…dzn⋅JdB=dx_1\,dy_1\ldots dz_n\cdot J

d B equals d x 1, d y 1, and so on up to d z n, times J.

where the factor J does not depend on the box sizes, nor on V∗V^*, that is, on where the wall is. J does not depend on where the boxes are either. Take a second set of boxes, of the same sizes, in other places inside V∗V^*; its part of B is dB′dB' with a factor J′J'. Since the sizes are equal,

dBdB′=JJ′\frac{dB}{dB'}=\frac{J}{J'}

d B over d B prime equals J over J prime.

Einstein's earlier theory gives these parts a meaning: dB/B is the probability that, at a moment chosen at random, the centres are in the given boxes. If the particles move independently of one another, to a sufficient approximation, the liquid is homogeneous and no forces act on the particles, then equal boxes are equally probable wherever they are, so

dBB=dB′B\frac{dB}{B}=\frac{dB'}{B}

d B over B equals d B prime over B.

and with the previous equation, J=J′J = J'.

Return to why equal boxes are equally probable. (included in this file)

So J depends neither on V∗V^* nor on where the particles are. Integrating over all positions of the n centres, each ranging over the volume V∗V^*, gives

B=∫J dx1…dzn=JV∗nB=\int J\,dx_1\ldots dz_n=JV^{*n}

B equals the integral of J over d x 1 through d z n, which equals J times V star to the power n.

and so the free energy is

F=−RTN{lg⁡J+nlg⁡V∗}F=-\frac{RT}{N}\left\{\lg J+n\lg V^*\right\}

F equals minus R T over N, times the logarithm of J plus n times the logarithm of V star.

The pressure on the wall is minus the rate at which F changes as V∗V^* grows:

p=−∂F∂V∗=RTV∗nN=RTN νp=-\frac{\partial F}{\partial V^*}=\frac{RT}{V^*}\frac{n}{N}=\frac{RT}{N}\,\nu

p equals minus the partial derivative of F with respect to V star, which equals R T over V star times n over N, which equals R T over N times nu.

Return to the pressure as the change of free energy with volume. (included in this file)

This shows, Einstein concludes, that osmotic pressure is a consequence of the molecular-kinetic theory of heat, and that on this theory equal numbers of dissolved molecules and suspended bodies behave exactly alike as regards osmotic pressure at great dilution. The question of how the theory avoids solving every molecular motion has a plain answer: the hard part of B, the factor J, is never computed. Only its independence of V∗V^* is needed, and that follows from equal boxes being equally probable. The volume enters only through V∗nV^{*n}, and its logarithm, nlg⁡V∗n\lg V^*, gives the pressure.

Assumptions

  • Einstein's statistical theory of heat of 1902 and 1903: for a system whose state variables change by equations whose rates satisfy the sum condition of §2, the entropy and the free energy are given by the logarithm of an integral over all its states.
  • The particles move independently of one another to a sufficient approximation, the liquid is homogeneous, and no forces act on the particles.
  • The total volume of the particles is small compared with the volume V* that holds them.

Limits

  • The footnote to the heading says this section, and Einstein's earlier papers on the foundations of thermodynamics, are not needed to understand the paper's results. The passage explains the section for the reader who wants to know how the law is obtained.
  • The argument finds only how the free energy depends on V*; it computes nothing else about the integral B.
  • Independence, homogeneity and the absence of forces are assumptions. With particles that act on one another, or crowded ones, J would depend on their positions and the pressure would depart from the dilute law.

Foundations included in this file

Reading a graph

How does a curve show the relationship between two physical quantities?

Walk for a minute and note how far you have gone every ten seconds: 0 m at the start, 14 m after ten seconds, 28 m after twenty, and so on. Draw a line along the bottom of a page for time and one up the side for distance, and put a dot for each pair. That is a graph: each dot is one moment, read across for the time and up for the distance.

The dots of a steady walk fall on a straight line, and the steeper the line, the faster the walk. This one climbs 14 metres for every 10 seconds: 1.4 metres per second. A flat line would mean standing still, and a line that bends upward would mean speeding up.

The papers draw other pairs the same way. Along the bottom goes what you choose or wait for, such as a position or a time; up the side goes what you then find, such as how many particles sit there. A bell-shaped curve, high in the middle and falling away on both sides, says that most particles are near the centre and fewer are farther out.

Worked example: Four particles that average to zero

  1. Four particles end at −3, −1, +1 and +3 micrometres from where they started.
  2. Mark those positions along the bottom and draw a bar one particle tall above each.
  3. The bars stand two on each side of zero, at the same distances out. The picture is symmetric, so the average position is 0.
  4. Yet no bar stands at zero: every particle moved. The graph shows at a glance what the average hides.

Where this lesson stops

Before reading a curve, read its two labels and their units, and ask whether its height is a count, a density or a running total.

This lesson builds on

Return to the chapter

A sign records direction

How can movement add up to zero?

Put the starting point at zero on a ruler. A final mark three units to the right has displacement +3; one three units to the left has displacement −3. Adding the signed displacements gives zero. Adding the distances from the start gives six.

Displacement compares where something ends with where it started, whatever it did in between. A walker who goes three units right and comes back has displacement zero after walking six.

Worked example: Two walkers, opposite directions, average zero

Two walkers both finish one unit from their start, one at −1 and the other at +1. The signed average is zero. The average distance is one.

Where this lesson stops

The sign names a direction relative to a chosen axis; it does not mean a negative distance.

This lesson builds on

Return to the chapter

Adding and averaging

What does an average keep, and what does it lose?

Four jars hold 3, 1, 1 and 3 marbles. Pour them together and there are 8. Share the 8 equally between the four jars and each holds 2. Two is the average, although no jar held 2 to begin with.

The average keeps the total and the count and nothing else, so 3, 1, 1, 3 and 2, 2, 2, 2 have the same average. Listing every value twice doubles both the total and the count, and the average stays at 2.

Worked example: Adding four numbers and dividing by four

  1. Add 3 + 1 + 1 + 3 to get 8.
  2. Count the values: there are four.
  3. Divide 8 by 4 to get 2.
  4. List the values twice: 16 divided by 8 is still 2.

Where this lesson stops

An average is the total shared equally among the observations, and two very different lists can have the same one.

Return to the chapter

Rates of change and derivatives

How can a body have a speed at a single instant?

A speed needs two readings of position and the time between them: divide the distance covered by the time taken, and you have the average speed over that interval. An instant has no duration, so that recipe cannot be used at an instant as it stands.

So shorten the interval and watch what happens. A ball rolls so that after t seconds it has gone 3t² metres: 3 m after one second, 12 m after two. Start at the one-second mark and average over shorter and shorter stretches:

  1. Over the next second it goes from 3 m to 12 m: 9 m in 1 s, an average of 9 m/s.
  2. Over the next tenth of a second it reaches 3.63 m: 0.63 m in 0.1 s, an average of 6.3 m/s.
  3. Over the next hundredth it reaches 3.0603 m: 0.0603 m in 0.01 s, an average of 6.03 m/s.
  4. Over the next thousandth it reaches 3.006003 m: 0.006003 m in 0.001 s, an average of 6.003 m/s.

The averages close in on 6 m/s. That number is the derivative: the ball's speed at the instant t = 1 s. The interval is never set to zero, which would mean dividing zero distance by zero time. It only shrinks, and the averages settle.

In symbols, with Δt for the short interval, the same recipe reads:

dxdt=lim⁡Δt→0x(t+Δt)−x(t)Δt\frac{dx}{dt} = \lim_{\Delta t \to 0} \frac{x(t+\Delta t) - x(t)}{\Delta t}

d x by d t is the value that the change in position divided by the change in time settles on as delta t shrinks towards zero.

A derivative carries units: those of the quantity that changes, divided by those of the quantity you vary. Position in metres, varied over time in seconds, gives metres per second. On a graph of position against time, the derivative at a point is the steepness of the curve there.

Einstein's Brownian paper needs rates like this for a concentration of particles, which changes along a tube and in time at once. With two things changing, a rate has to say which one it follows, so the paper writes a curly ∂ and holds the other fixed:

∂f∂t=D ∂2f∂x2\frac{\partial f}{\partial t} = D\,\frac{\partial^2 f}{\partial x^2}

The partial derivative of f with respect to t equals D times the second partial derivative of f with respect to x.

The lesson on partial derivatives, listed at the end of this page, reads that equation one piece at a time.

Worked example: Why the averages settle on 6 m/s

The list of averages can be explained, not only watched. Start at any time t and let the interval be Δt. In that interval the ball moves 3(t + Δt)² − 3t² metres. Since (t + Δt)² is t² + 2tΔt + (Δt)², that distance is:

3(t+Δt)2−3t2=6t Δt+3(Δt)23(t+\Delta t)^2 - 3t^2 = 6t\,\Delta t + 3(\Delta t)^2

Three times t plus delta t, squared, minus three t squared, equals six t delta t plus three delta t squared.

Divide that distance by the time Δt to get the average speed over the interval:

6t Δt+3(Δt)2Δt=6t+3 Δt\frac{6t\,\Delta t + 3(\Delta t)^2}{\Delta t} = 6t + 3\,\Delta t

Six t delta t plus three delta t squared, divided by delta t, equals six t plus three delta t.

The 3Δt term shrinks with the interval and leaves 6t metres per second. At t = 1 s that is 6 m/s. The extra 3, 0.3, 0.03 and 0.003 in the list above are that 3Δt term, for Δt of 1, 0.1, 0.01 and 0.001 seconds.

Where this lesson stops

A derivative is the value an average rate settles on as the interval shrinks, in units of one quantity per unit of the other. The interval shrinks but is never set to zero, so nothing is ever divided by zero.

This lesson builds on

Return to the chapter

Entropy and the number of ways

Why is entropy the logarithm of a probability?

Four particles move independently in a box. The chance that all four happen to be in the left half at a given moment is ½ × ½ × ½ × ½ = 1/16. For a hundred particles it is (½)¹⁰⁰, about 8 × 10⁻³¹: possible, but never seen. A state is more probable when more of the ways the particles can be arranged produce it.

Entropy measures this probability on a logarithmic scale. Boltzmann's principle, as §5 of the light-quanta paper states it, makes the entropy of a system a function of the probability W of its state, and that function must be a logarithm:

S−S0=kBln⁡WS - S_0 = k_B \ln W

S minus S nought equals k B times the natural logarithm of W.

Why a logarithm: two independent systems have a combined probability W = W₁ × W₂, while their entropies add, S = S₁ + S₂. Only a logarithm turns the product into the sum, and that is the argument §5 gives. The constant kBk_B is R/N, which is how the paper writes it, and the paper prints the natural logarithm as lg.

Apply it to the box. Squeezing n independent particles from a volume v₀ into a part v of it has probability W = (v/v₀)ⁿ, so the entropy changes by S − S₀ = n kBk_B ln(v/v₀), a decrease, since v/v₀ is less than 1. §5 derives exactly this, and remarks that it needs no assumption about the law by which the molecules move.

The probability here is a frequency: the fraction of moments at which the state is found. §5 insists on this, and criticizes calculations in which the equally probable cases are simply postulated.

Worked example: Squeezing a gas into half its volume

  1. Take one gram-molecule of gas, n = N = 6.02 × 10²³ particles, and ask for all of them in the left half: v/v₀ = ½.
  2. The probability is (½) to the power 6.02 × 10²³, far too small to write out. Its logarithm is N ln ½, about −4.2 × 10²³.
  3. The entropy change is kBk_B × N × ln ½ = R ln ½ = 8.314 × (−0.693), about −5.76 joules per kelvin.
  4. Thermodynamics gives the same number for compressing an ideal gas to half its volume at fixed temperature, R ln ½. Counting ways and measuring heat agree.

§6 of the paper finds the same form for dilute radiation, with E/(βν) standing where the gas has Rn/N, and reads it as radiation behaving, in this respect, like independent quanta.

Where this lesson stops

This lesson stops at entropy as the logarithm of a probability. What the comparison with radiation shows, and what it does not, is §6 of the light-quanta paper.

This lesson builds on

Return to the chapter

Functions and graphs

How does a graph record the way one quantity depends on another?

A bath fills from a tap. After one minute it holds 10 litres, after two 20, after three 30. For each time there is exactly one amount of water, and that pairing, one output for each input, is a function. Plot the pairs with time along the bottom and litres up the side and they fall on a straight line that climbs 10 litres for every minute. The slope, in litres per minute, is how fast the bath fills.

In the papers the pairs are physical: the mean square displacement of a Brownian particle for each elapsed time, or the concentration for each position along a tube. A function says how one thing responds when another changes.

A graph puts the input along the horizontal axis and the output up the vertical one, so each point is one pair. Both axes carry units, and so does everything read off the graph: the slope of mean square displacement against time is in square micrometres per second, and the area under a probability density is a plain probability.

A formula such as ⟨x²⟩ = 2Dt describes an average over many particles. One particle's squared displacement scatters widely around the line; the average over many lies close to it. The Brownian paper predicts the average, and ends by hoping that a researcher will soon decide the question by observation.

Worked example: Mean square displacement plotted against time

⟨x2⟩=2Dt\langle x^2\rangle = 2Dt

The mean square displacement equals two times the diffusion coefficient times the time.

  1. Take D = 0.5 square micrometres per second, so 2D = 1 square micrometre per second.
  2. At t = 1 s the mean square displacement is 1 µm²; at t = 4 s it is 4 µm²; at t = 9 s, 9 µm².
  3. Plotted against time, these points lie on a straight line through the origin whose slope is 2D, 1 µm² per second.
  4. Their square roots, the root-mean-square displacements, are 1, 2 and 3 µm: four times the time gives twice the distance.

Where this lesson stops

A graph is a record of paired quantities, each with its units. A slope read off it has units too, and a slope stated without them explains nothing.

This lesson builds on

Return to the chapter

Logarithms: turning products into sums

Why does entropy need a logarithm?

Adding is easier than multiplying, and a logarithm turns one into the other. The natural logarithm of 2 is about 0.693, and of 4 about 1.386. Add them and you get 2.079, which is the natural logarithm of 8, the product of 2 and 4.

The natural logarithm ln x answers one question: to what power must e, about 2.718, be raised to give x? Raised to 0.693, e gives 2, and raised to 1.386 it gives 4. Multiplying two powers of e adds the powers, so raised to 2.079 it gives 8. Written for any numbers, the logarithm of a product is the sum of the logarithms, and a power becomes a multiple:

ln⁡(AB)=ln⁡A+ln⁡B,ln⁡(f n)=nln⁡f\ln(AB) = \ln A + \ln B, \qquad \ln\left(f^{\,n}\right) = n \ln f

The log of A times B is log A plus log B, and the log of f to the power n is n times log f.

Entropy needs exactly this. In §5 of the light-quanta paper, Einstein takes two independent systems. Their entropies add, S = S₁ + S₂, while the probabilities of their states multiply, W = W₁W₂. If entropy is some function φ of probability, then φ(W₁W₂) = φ(W₁) + φ(W₂), and he concludes that φ is a logarithm:

S−S0=RNln⁡WS - S_0 = \frac{R}{N} \ln W

S minus S nought equals R over N times the natural log of W.

R is the gas constant and N the number of molecules in a mole, so R/N is what is now called Boltzmann's constant. The step from φ(W₁W₂) = φ(W₁) + φ(W₂) to the logarithm needs φ to be a smooth function, which Einstein takes for granted.

Einstein prints lg for the natural logarithm, as German physics did in 1905. Today lg usually means the base-10 logarithm and ln the natural one. So a printed lg 2 means ln 2, about 0.693, not log₁₀ 2, about 0.301.

Worked example: Ten molecules found in half the volume

In §5 Einstein asks how probable it is that n molecules, each moving independently through a volume v₀, are all found at one moment in a part v of it. Each molecule is in that part with probability v/v₀, and independent probabilities multiply:

W=(vv0)nW = \left(\frac{v}{v_0}\right)^{n}

W equals v over v nought, to the power n.

  1. Take n = 10 molecules and v = v₀/2. Then W = (1/2)¹⁰ = 1/1024, about 0.000977.
  2. The logarithm turns the tenth power into ten times: ln W = 10 × ln(1/2) = 10 × (−0.693) = −6.93.
  3. So the entropy of the state with every molecule in one half is lower by 6.93 × R/N: 6.93 times Boltzmann's constant.

For a mole of gas, n = N, and the same formula gives S − S₀ = R ln(v/v₀). That is the entropy change thermodynamics gives for an ideal gas compressed from v₀ to v at constant temperature.

Where this lesson stops

A logarithm turns products into sums, which is why it connects probabilities that multiply to entropies that add. In the 1905 papers, lg means the natural logarithm.

This lesson builds on

Return to the chapter

Partial derivatives and held-fixed quantities

When a quantity depends on two things, what does its rate of change mean?

The concentration of particles in a tube depends on where you look, x, and when you look, t. Ask how fast the concentration changes and there are two different answers: how it changes as you move along the tube at one instant, and how it changes at one place as time passes.

A partial derivative picks one of them. ∂f/∂x is the rate of change along the tube with the time t held fixed, like comparing points in one photograph. ∂f/∂t is the rate of change in time at a fixed position x, like watching one spot under the microscope. The curly ∂ says that something else is being held fixed, and a subscript names it:

(∂f∂x)t(∂f∂t)x\left(\frac{\partial f}{\partial x}\right)_t \qquad \left(\frac{\partial f}{\partial t}\right)_x

The partial derivative of f with respect to x at fixed t, and the partial derivative of f with respect to t at fixed x.

Thermodynamics depends on the same care. How much a gas's volume changes as the pressure changes depends on whether its temperature T is held fixed, with heat flowing in and out, or its entropy S is held fixed, with the gas insulated. The two derivatives have different values, and ∂V/∂p written alone does not say which is meant:

(∂V∂p)Tversus(∂V∂p)S\left(\frac{\partial V}{\partial p}\right)_T \quad \text{versus} \quad \left(\frac{\partial V}{\partial p}\right)_S

The partial derivative of V with respect to p at fixed temperature T, versus the same derivative at fixed entropy S.

In §3 of the light-quanta paper, radiation fills a fixed volume and φ is its entropy per unit volume and per unit interval of frequency, a function of the energy density ρ and of ν. Varying ρ with the frequency ν held fixed, Einstein finds that the derivative is 1/T:

(∂φ∂ρ)ν=1T\left(\frac{\partial \varphi}{\partial \rho}\right)_{\nu} = \frac{1}{T}

The partial derivative of phi with respect to rho, at fixed frequency nu, equals one over T.

In §6 of the relativity paper, Maxwell's equations are rewritten in the coordinates ξ, η, ζ, τ of a moving system. A derivative with respect to x taken with y, z and t held fixed then becomes a combination of a derivative with respect to ξ and one with respect to τ, because holding the resting time t fixed does not hold the moving time τ fixed.

Worked example: The same diffusion profile, differentiated two ways

Take the spreading profile of the Brownian paper's §4, written for a single particle:

f(x,t)=14πDt e−x2/4Dtf(x,t) = \frac{1}{\sqrt{4\pi D t}}\,e^{-x^2/4Dt}

f of x and t equals one over the square root of four pi D t, times e to the minus x squared over four D t.

Hold t fixed and differentiate with respect to x. Only the exponent depends on x, and the derivative of −x²/4Dt with respect to x is −x/2Dt:

(∂f∂x)t=−x2Dt f(x,t)\left(\frac{\partial f}{\partial x}\right)_t = -\frac{x}{2Dt}\,f(x,t)

The partial derivative of f with respect to x at fixed t equals minus x over two D t, times f.

That is the slope of the profile in one snapshot: downhill to the right of the centre, uphill to the left. Now hold x fixed and differentiate with respect to t. Both the factor in front and the exponent depend on t:

(∂f∂t)x=(−12t+x24Dt2)f(x,t)\left(\frac{\partial f}{\partial t}\right)_x = \left(-\frac{1}{2t} + \frac{x^2}{4Dt^2}\right) f(x,t)

The partial derivative of f with respect to t at fixed x equals minus one over two t, plus x squared over four D t squared, all times f.

That is what a probe fixed at one place records. Near the centre, where x² is less than 2Dt, the bracket is negative: the concentration falls as the cloud spreads away. Farther out it is positive: the concentration rises as the cloud arrives. Differentiate the first result once more with respect to x and multiply by D, and you get the second. That is the §4 equation ∂f/∂t = D ∂²f/∂x², a rate at one place equal to a curvature at one time.

Where this lesson stops

A partial derivative is not defined until you have said which quantities are held fixed.

This lesson builds on

Return to the chapter

Probability and independence

When you square a sum of random steps, why do the mixed terms average to zero?

Flip a coin twice and take one step for each flip: a metre to the right for heads, a metre to the left for tails. How far from the start do you end up, on average?

  1. The four outcomes, right-right, right-left, left-right and left-left, are equally likely: one chance in four each.
  2. They leave you 2 m to the right, back at the start, back at the start, or 2 m to the left.
  3. Square those distances so that left and right count alike: 4, 0, 0 and 4. Their average is 2, one square metre for each step.

Combining the two steps added nothing beyond one square metre each. In symbols, with Δ₁ and Δ₂ for the two steps, squaring the sum gives each step squared and a mixed term, twice their product:

(Δ1+Δ2)2=Δ12+2Δ1Δ2+Δ22(\Delta_1 + \Delta_2)^2 = \Delta_1^2 + 2\Delta_1\Delta_2 + \Delta_2^2

Delta one plus delta two, squared, equals delta one squared plus two delta one delta two plus delta two squared.

In the coin walk the product of the two steps is +1 for right-right and left-left and −1 for the other two, so on average it is zero and the mixed term adds nothing. Two conditions make that happen. Independence: learning the first step does not change the chances for the second, so the chance of a pair is the product of the two chances. Centring: each step averages zero, left as likely as right. Then the average of the product is the product of the averages, zero times zero:

⟨Δ1Δ2⟩=⟨Δ1⟩⟨Δ2⟩=0\langle \Delta_1 \Delta_2 \rangle = \langle \Delta_1 \rangle \langle \Delta_2 \rangle = 0

The average of delta one times delta two equals the average of delta one times the average of delta two, which is zero.

Both conditions matter. If each step drifts, averaging +1, the product of the averages is 1, not 0. If the second step tends to copy the first, the product is positive more often than negative. The Brownian paper's §4 assumes that a particle's motions in successive intervals are independent, as long as the intervals are not chosen too small.

Worked example: What changes when the second step copies the first

  1. Now let the second flip always copy the first. Only right-right and left-left can happen, one chance in two each.
  2. The product of the two steps is +1 every time, so the mixed term no longer averages zero.
  3. You always end 2 m from the start, so the squared distance is 4 every time: the average is 4, not 2. Independence removed that extra 2.

Where this lesson stops

Before dropping a mixed term, ask two things: does learning one step change the chances for the next, and does each step average zero?

This lesson builds on

Return to the chapter

Static laboratory examples

bm-03

This laboratory is not included in this file. The chapter's authored foundation examples remain available above. Open bm-03 online

Source references

  1. A. Einstein, Über einen die Erzeugung und Verwandlung des Lichtes betreffenden heuristischen Gesichtspunkt. Annalen der Physik (4), 17, 132–148 (1905).
  2. A. Einstein, On the motion of particles suspended in liquids at rest required by the molecular-kinetic theory of heat. Annalen der Physik (4), 17, 549–560 (1905), §§1–5. Bibliographic pointer; this preview is not a source transcription or translation.