Brownian motion · Learning from a finite sample

From wandering
to a number.

An equation predicts how particles spread. Turn the question around: what can a finite set of displacements tell you about its hidden parameters, and what must you know independently?

Read the inference argument →

Bring your own calibrated trajectory CSV →

BM-07 · The inverse problem

What can wandering reveal?

Static worked example

These positions were generated with a hidden molecular number. First ask what the observations identify. Then declare the missing information, estimate the number, and test what a confidence interval does across hypothetical repeats.

Observation set

The synthetic set checks inference machinery on data made with a hidden number. It is not evidence that molecules exist. Perrin 1909 waits on the admitted HistoricalDataset; this instrument will not invent table numbers. Kitchen CSV is analyzed by the existing kitchen session; BM-07 consumes that session and does not re-parse video.

1 · Observe the same path

The fixed recording lasts 1024 seconds. Spacing, coordinate count, sample count and estimator changes reuse it. No overlapping windows, localization noise, exposure blur or censoring are admitted by these intervals.

2 · Declare inference inputs

These are inference assumptions, not generator controls. Changing them never moves a recorded position. A radius derived from these same displacements using an assumed molecular number would be circular, not independent information.

Declare input uncertainty

Each input interval is its stated value plus or minus the relative bound below. Its declared coverage must describe a valid measurement procedure. Bounds alone are not confidence intervals. Calibration, timing and the chosen gas constant remain exact in this exercise.

The combined procedure allocates the remaining error probability to the diffusion interval and uses a union bound. It requires no independence between the three input intervals, but cannot create a coverage guarantee from undeclared coverage.

3 · Apply the request
Change the synthetic generator

Changing these fields starts a new physical run. The hidden molecular number is drawn on a separate stream from 3 × 10²³ to 1.2 × 10²⁴ mol⁻¹. The generator uses the explicitly chosen R = 8.314471 J/(mol K), not a modern Boltzmann constant that would preselect the answer.

These comparison buttons use accepted settings, not draft edits. “Use declared generator conditions” supplies the known setup of this synthetic exercise; it is not a real independent measurement.

The CSV contains every selected position and displacement in SI units with generator metadata. It contains synthetic data, not observations of a real suspension. Opening a shared link starts no worker.

Static worked example. No calculation has been started in this browser.

Accepted seed 1905; 50 non-overlapping displacements; 2 coordinates; spacing 1 s. Known zero drift · unbiased. Assumed T = 293.15 K, η = 1 mPa s; independent radius not declared.

7.90940-7.909402550Displacement (μm)Observation time (s)
One synthetic path, observed at the accepted spacing. Solid: x; dashed: y when selected. Connecting lines join observations; they do not claim a resolved microscopic trajectory.

What the data identify

Accepted estimate and conditional diffusion uncertainty
Estimated D (μm²/s)0.88271
95% diffusion interval (μm²/s)0.681311.1893
Degrees of freedom100
Compatible radius × number (m/mol)1.4649e17

A family before a single number

At the assumed temperature and viscosity, the estimated diffusion scale constrains a product: radius times molecular number. A larger radius and a smaller number can fit exactly the same point estimate.

0.114.6490.447213.275620.73244N (10²³ mol⁻¹); logarithmic axesAssumed radius (μm)
Each point on the solid curve gives the same diffusion estimate at the stated temperature and viscosity. This is a compatible family, not a confidence region and not a second measurement of the radius.
Read all compatible pairs
Radius–number family for this accepted estimate
Radius (μm)N (10²³ mol⁻¹)
0.114.649
0.1077813.592
0.1161612.611
0.1251911.701
0.1349310.857
0.1454210.073
0.156739.3465
0.168928.6721
0.182068.0463
0.196217.4657
0.211476.927
0.227926.4272
0.245655.9634
0.264755.5331
0.285345.1339
0.307534.7634
0.331454.4197
0.357224.1008
0.3853.8049
0.414943.5303
0.447213.2756
0.481993.0392
0.519482.8199
0.559882.6165
0.603422.4277
0.650342.2525
0.700922.0899
0.755431.9391
0.814181.7992
0.87751.6694
0.945741.5489
1.01931.4372
1.09861.3335
1.1841.2372
1.27611.148
1.37531.0651
1.48230.98828
1.59750.91696
1.72180.8508
1.85570.78941
20.73244

This checks the inference method on data made with a hidden number. It is not evidence that molecules exist.

Condition on the missing information

Conditional on exact input values; target 95%
Recovered N (10²³ mol⁻¹)The data constrain a radius–number product, not radius and molecular number separately. Declare an independent radius to condition the estimate.
Selected interval (10²³ mol⁻¹)The data constrain a radius–number product, not radius and molecular number separately. Declare an independent radius to condition the estimate.

This is recovery of a synthetic generating parameter, not a measurement of the modern Avogadro constant. The inverted bounds exchange endpoints. A finite confidence interval is not a posterior probability assigned to this one fixed parameter.

Estimator bias and small-sample limits

The known-zero-drift estimate uses every coordinate without subtracting a fitted mean. Fitting a separate drift in each coordinate consumes degrees of freedom. The centered maximum-likelihood estimate divides by the original sample count; the unbiased version uses one fewer. Both use the same correctly rescaled confidence interval.

Expected ratios under the ideal model with correct input values
Mean estimated D / true D1
Mean recovered N / true N1.0204
Variance of recovered N / true N0.021692

An unbiased diffusion estimate does not have an unbiased reciprocal. With too few degrees of freedom, the reciprocal’s mean or variance does not exist even though a particular trial can produce a finite estimate. With one displacement and fitted drift, spread is underdetermined.

The hidden answer is a learning device, not a secret: the generated browser data can be inspected.

Inspect the accepted observations
All selected synthetic positions and increments; μm and seconds
Timex (μm)y (μm)Δx (μm)Δy (μm)
000
1-2.3879-0.70732-2.3879-0.70732
2-1.6243-1.0740.76359-0.36669
3-2.41131.5563-0.787032.6303
4-0.774542.92971.63681.3734
51.8554.35882.62951.4292
60.696954.734-1.1580.3752
7-0.739544.6997-1.4365-0.034373
82.12014.31452.8597-0.38513
90.94444.4184-1.17570.10385
100.550875.964-0.393541.5456
110.636027.90940.0851571.9454
120.924166.38230.28814-1.5271
132.37567.07661.45150.69424
142.63467.38120.2590.30462
152.70687.89370.0721210.51253
161.6457.0356-1.0618-0.85806
171.50497.7346-0.140090.69891
180.00875047.6172-1.4962-0.11731
191.68666.26461.6779-1.3527
200.639126.1636-1.0475-0.10093
21-0.665025.5842-1.3041-0.57943
22-2.91346.1162-2.24840.532
23-2.48725.71310.42622-0.40314
24-2.55997.2136-0.0726521.5005
25-1.93646.54180.62345-0.67176
26-3.38136.6304-1.44490.088609
27-2.33245.58091.0489-1.0495
28-2.04592.95560.28654-2.6253
29-3.56293.7742-1.5170.81859
30-0.164843.15753.3981-0.61668
31-1.45193.1915-1.28710.033954
32-2.06791.4143-0.61602-1.7772
33-2.51491.6097-0.446960.19549
34-1.68314.78810.831763.1783
35-0.533073.80341.1501-0.98462
36-1.64685.9611-1.11382.1577
37-1.25314.44180.39373-1.5193
38-2.23365.296-0.98050.85426
39-1.74457.05820.489111.7621
40-1.48247.33320.262110.275
410.977316.6242.4597-0.70916
422.67697.04391.69960.41987
430.470835.5085-2.2061-1.5354
440.4184.7481-0.05283-0.76038
450.919452.19810.50145-2.55
463.82790.855282.9085-1.3428
473.348-0.76496-0.47992-1.6202
482.4238-0.74306-0.92420.021895
490.59245-1.1868-1.8314-0.44372
500.0198980.25574-0.572551.4425

What does coverage mean?

Repeat the observation procedure on other hypothetical paths with the same generating parameter. Each interval moves; the parameter does not. Change an inference assumption while keeping those paths to see how a wrong radius can spoil molecular-number coverage even when diffusion intervals behave well.

Uses the accepted settings. No seed search, discarded failures or redraws to reach the target fraction. In combined-input mode this view is unavailable: no repeated input-measurement procedure has been specified. After a form edit, run coverage explicitly again.

Diffusivity procedure

Run the hypothetical experiments explicitly to inspect interval coverage.

Molecular-number procedure

Run the hypothetical experiments explicitly to inspect interval coverage.

Same scientific action without the plot

Choose the observation set, estimator, and interval kind from the lists. Type the displacement count M and the inference temperature, viscosity, and radius. Read the table of estimated D and its interval, then N and its interval with the wording above, the inverse-bias note, and the identifiability family. The plots are a view of that same accepted snapshot.

Not modeled: localization error, blur, correlated or irregularly timed increments, and censoring (see BM-08 and kitchen mode); non-Gaussian increments; time-varying drift; polydispersity within one track set; wall effects; uncertainty in C without declared coverages; uncertainty in the gas constant itself.

Accepted calculation identity and limits

Run bm07-_R_6av5uiv5b_/run/1; snapshot 1; revisions {"input":1,"observer":0,"measurement":0,"estimator":0}. Host reference owner inference.bm07, scenario scenario-bm07-hidden-number.

Primary recording draws: 16385. Work for this request: 16385 draws, including 0 for hypothetical experiments. Retained primary recording: 65552 bytes. Reused worker recording: no.

Source identity: source:sha256:2b658a4b9461d836ec8ed1abf378bf46cab28ae0990ee15487493a19446add28. This host preview has no historical data importer or camera-noise fit. Browser arithmetic is not a claim of strict cross-engine WASM replay.

Each placement has separate settings and worker ownership. The same seed initially reproduces the same synthetic data; choose a different seed for an independent trial.

Open the inverse argument

The same positions, different questions

For independent, equally spaced Gaussian increments in d coordinates with known zero drift, average their squared lengths and divide by twice the number of coordinates and the time interval. The result estimates the diffusion coefficient.

D^=12dMΔti=1MΔri2,q=dM\widehat D=\frac{1}{2dM\Delta t}\sum_{i=1}^{M}\lVert\Delta\mathbf r_i\rVert^2,\qquad q=dM

When drift is fitted from the same observations, subtract the mean increment in each coordinate. The unbiased spread estimate uses M − 1 in place of M and q = d(M − 1). The centered maximum-likelihood estimate instead divides by M and must be rescaled before applying the same interval procedure.

An interval for a procedure

Under this ideal model, q times the unbiased diffusion estimate divided by the true diffusivity follows a chi-square distribution. Its quantiles provide the following conditional interval. More independent observations generally tighten uncertainty; changing the assumptions changes what the interval means.

[qD^unbiasedχq,1α/22, qD^unbiasedχq,α/22]\left[\frac{q\widehat D_{\mathrm{unbiased}}}{\chi^2_{q,1-\alpha/2}},\ \frac{q\widehat D_{\mathrm{unbiased}}}{\chi^2_{q,\alpha/2}}\right]

At 95% coverage, the procedure covers the fixed true value in 95% of repetitions under its assumptions. That is not a 95% posterior probability for one parameter in one realized interval. The repeated-trial plots retain the misses as well as the successes.

Why an independent radius matters

Na=RT6πηD,N^=RT6πηaD^N a=\frac{RT}{6\pi\eta D},\qquad \widehat N=\frac{RT}{6\pi\eta a\widehat D}

With temperature and viscosity given, displacements identify the product Na. They cannot separately identify N and the radius a. Declaring a radius permits a conditional inversion, but deriving that radius from the same data using an assumed N would merely return the assumption.

Inversion also reverses the interval’s endpoints. It does not preserve unbiasedness: for the known-zero-drift estimate, the mean recovered-to-true number ratio is q/(q − 2) when q is greater than two. A point estimate can be finite even when that repeated-estimate mean does not exist.

Do not hide input uncertainty

The conditional interval treats the stated inputs as exact. The combined procedure uses the declared temperature, viscosity and radius intervals and adds their error probabilities with the diffusion interval’s error probability. This union-bound guarantee is conservative and does not require their statistical independence. It remains conditional on exact calibration, timing and the chosen gas constant in this preview.

Worked interval (readable without JavaScript)

d = 2, M = 50, q = 100, D̂ = 0.42944 μm²/s, 95% chi-square interval. Not a molecular count.
QuantityValue
Diffusion interval[0.331457, 0.578589] μm²/s
Modern-SI inversion (consistency check)N̂ = 6.02213 × 10²³ mol⁻¹
Inverse-bias factor at q = 100q/(q − 2) = 1.020408; an unbiased D̂ is not unbiased after inversion

A synthetic recovery is not a new molecular count

This exercise draws a hidden number and generates data from it using a chosen gas constant. Recovering that number tests inference under the model; it does not measure a real suspension. With modern SI constants, the Avogadro constant is defined, so an observational inversion would be a consistency check rather than an independent determination. Reviewed historical observations and independent gas-constant measurements are not supplied here. This instrument uses ideal positions; the separate camera laboratory adds an explicitly later measurement model.

NIST: chi-square confidence limits for a normal variance · BIPM: SI units and defining constants

Next: keep the particle, change the camera, and test the inference →