Brownian motion · Learning from a finite sample
From wandering
to a number.
An equation predicts how particles spread. Turn the question around: what can a finite set of displacements tell you about its hidden parameters, and what must you know independently?
BM-07 · The inverse problem
What can wandering reveal?
These positions were generated with a hidden molecular number. First ask what the observations identify. Then declare the missing information, estimate the number, and test what a confidence interval does across hypothetical repeats.
These comparison buttons use accepted settings, not draft edits. “Use declared generator conditions” supplies the known setup of this synthetic exercise; it is not a real independent measurement.
The CSV contains every selected position and displacement in SI units with generator metadata. It contains synthetic data, not observations of a real suspension. Opening a shared link starts no worker.
Static worked example. No calculation has been started in this browser.
Accepted seed 1905; 50 non-overlapping displacements; 2 coordinates; spacing 1 s. Known zero drift · unbiased. Assumed T = 293.15 K, η = 1 mPa s; independent radius not declared.
What the data identify
| Estimated D (μm²/s) | 0.88271 |
|---|---|
| 95% diffusion interval (μm²/s) | 0.68131 – 1.1893 |
| Degrees of freedom | 100 |
| Compatible radius × number (m/mol) | 1.4649e17 |
A family before a single number
At the assumed temperature and viscosity, the estimated diffusion scale constrains a product: radius times molecular number. A larger radius and a smaller number can fit exactly the same point estimate.
Read all compatible pairs
| Radius (μm) | N (10²³ mol⁻¹) |
|---|---|
| 0.1 | 14.649 |
| 0.10778 | 13.592 |
| 0.11616 | 12.611 |
| 0.12519 | 11.701 |
| 0.13493 | 10.857 |
| 0.14542 | 10.073 |
| 0.15673 | 9.3465 |
| 0.16892 | 8.6721 |
| 0.18206 | 8.0463 |
| 0.19621 | 7.4657 |
| 0.21147 | 6.927 |
| 0.22792 | 6.4272 |
| 0.24565 | 5.9634 |
| 0.26475 | 5.5331 |
| 0.28534 | 5.1339 |
| 0.30753 | 4.7634 |
| 0.33145 | 4.4197 |
| 0.35722 | 4.1008 |
| 0.385 | 3.8049 |
| 0.41494 | 3.5303 |
| 0.44721 | 3.2756 |
| 0.48199 | 3.0392 |
| 0.51948 | 2.8199 |
| 0.55988 | 2.6165 |
| 0.60342 | 2.4277 |
| 0.65034 | 2.2525 |
| 0.70092 | 2.0899 |
| 0.75543 | 1.9391 |
| 0.81418 | 1.7992 |
| 0.8775 | 1.6694 |
| 0.94574 | 1.5489 |
| 1.0193 | 1.4372 |
| 1.0986 | 1.3335 |
| 1.184 | 1.2372 |
| 1.2761 | 1.148 |
| 1.3753 | 1.0651 |
| 1.4823 | 0.98828 |
| 1.5975 | 0.91696 |
| 1.7218 | 0.8508 |
| 1.8557 | 0.78941 |
| 2 | 0.73244 |
This checks the inference method on data made with a hidden number. It is not evidence that molecules exist.
Condition on the missing information
| Recovered N (10²³ mol⁻¹) | The data constrain a radius–number product, not radius and molecular number separately. Declare an independent radius to condition the estimate. |
|---|---|
| Selected interval (10²³ mol⁻¹) | The data constrain a radius–number product, not radius and molecular number separately. Declare an independent radius to condition the estimate. |
This is recovery of a synthetic generating parameter, not a measurement of the modern Avogadro constant. The inverted bounds exchange endpoints. A finite confidence interval is not a posterior probability assigned to this one fixed parameter.
Estimator bias and small-sample limits
The known-zero-drift estimate uses every coordinate without subtracting a fitted mean. Fitting a separate drift in each coordinate consumes degrees of freedom. The centered maximum-likelihood estimate divides by the original sample count; the unbiased version uses one fewer. Both use the same correctly rescaled confidence interval.
| Mean estimated D / true D | 1 |
|---|---|
| Mean recovered N / true N | 1.0204 |
| Variance of recovered N / true N | 0.021692 |
An unbiased diffusion estimate does not have an unbiased reciprocal. With too few degrees of freedom, the reciprocal’s mean or variance does not exist even though a particular trial can produce a finite estimate. With one displacement and fitted drift, spread is underdetermined.
The hidden answer is a learning device, not a secret: the generated browser data can be inspected.
Inspect the accepted observations
| Time | x (μm) | y (μm) | Δx (μm) | Δy (μm) |
|---|---|---|---|---|
| 0 | 0 | 0 | ||
| 1 | -2.3879 | -0.70732 | -2.3879 | -0.70732 |
| 2 | -1.6243 | -1.074 | 0.76359 | -0.36669 |
| 3 | -2.4113 | 1.5563 | -0.78703 | 2.6303 |
| 4 | -0.77454 | 2.9297 | 1.6368 | 1.3734 |
| 5 | 1.855 | 4.3588 | 2.6295 | 1.4292 |
| 6 | 0.69695 | 4.734 | -1.158 | 0.3752 |
| 7 | -0.73954 | 4.6997 | -1.4365 | -0.034373 |
| 8 | 2.1201 | 4.3145 | 2.8597 | -0.38513 |
| 9 | 0.9444 | 4.4184 | -1.1757 | 0.10385 |
| 10 | 0.55087 | 5.964 | -0.39354 | 1.5456 |
| 11 | 0.63602 | 7.9094 | 0.085157 | 1.9454 |
| 12 | 0.92416 | 6.3823 | 0.28814 | -1.5271 |
| 13 | 2.3756 | 7.0766 | 1.4515 | 0.69424 |
| 14 | 2.6346 | 7.3812 | 0.259 | 0.30462 |
| 15 | 2.7068 | 7.8937 | 0.072121 | 0.51253 |
| 16 | 1.645 | 7.0356 | -1.0618 | -0.85806 |
| 17 | 1.5049 | 7.7346 | -0.14009 | 0.69891 |
| 18 | 0.0087504 | 7.6172 | -1.4962 | -0.11731 |
| 19 | 1.6866 | 6.2646 | 1.6779 | -1.3527 |
| 20 | 0.63912 | 6.1636 | -1.0475 | -0.10093 |
| 21 | -0.66502 | 5.5842 | -1.3041 | -0.57943 |
| 22 | -2.9134 | 6.1162 | -2.2484 | 0.532 |
| 23 | -2.4872 | 5.7131 | 0.42622 | -0.40314 |
| 24 | -2.5599 | 7.2136 | -0.072652 | 1.5005 |
| 25 | -1.9364 | 6.5418 | 0.62345 | -0.67176 |
| 26 | -3.3813 | 6.6304 | -1.4449 | 0.088609 |
| 27 | -2.3324 | 5.5809 | 1.0489 | -1.0495 |
| 28 | -2.0459 | 2.9556 | 0.28654 | -2.6253 |
| 29 | -3.5629 | 3.7742 | -1.517 | 0.81859 |
| 30 | -0.16484 | 3.1575 | 3.3981 | -0.61668 |
| 31 | -1.4519 | 3.1915 | -1.2871 | 0.033954 |
| 32 | -2.0679 | 1.4143 | -0.61602 | -1.7772 |
| 33 | -2.5149 | 1.6097 | -0.44696 | 0.19549 |
| 34 | -1.6831 | 4.7881 | 0.83176 | 3.1783 |
| 35 | -0.53307 | 3.8034 | 1.1501 | -0.98462 |
| 36 | -1.6468 | 5.9611 | -1.1138 | 2.1577 |
| 37 | -1.2531 | 4.4418 | 0.39373 | -1.5193 |
| 38 | -2.2336 | 5.296 | -0.9805 | 0.85426 |
| 39 | -1.7445 | 7.0582 | 0.48911 | 1.7621 |
| 40 | -1.4824 | 7.3332 | 0.26211 | 0.275 |
| 41 | 0.97731 | 6.624 | 2.4597 | -0.70916 |
| 42 | 2.6769 | 7.0439 | 1.6996 | 0.41987 |
| 43 | 0.47083 | 5.5085 | -2.2061 | -1.5354 |
| 44 | 0.418 | 4.7481 | -0.05283 | -0.76038 |
| 45 | 0.91945 | 2.1981 | 0.50145 | -2.55 |
| 46 | 3.8279 | 0.85528 | 2.9085 | -1.3428 |
| 47 | 3.348 | -0.76496 | -0.47992 | -1.6202 |
| 48 | 2.4238 | -0.74306 | -0.9242 | 0.021895 |
| 49 | 0.59245 | -1.1868 | -1.8314 | -0.44372 |
| 50 | 0.019898 | 0.25574 | -0.57255 | 1.4425 |
What does coverage mean?
Repeat the observation procedure on other hypothetical paths with the same generating parameter. Each interval moves; the parameter does not. Change an inference assumption while keeping those paths to see how a wrong radius can spoil molecular-number coverage even when diffusion intervals behave well.
Uses the accepted settings. No seed search, discarded failures or redraws to reach the target fraction. In combined-input mode this view is unavailable: no repeated input-measurement procedure has been specified. After a form edit, run coverage explicitly again.
Diffusivity procedure
Run the hypothetical experiments explicitly to inspect interval coverage.
Molecular-number procedure
Run the hypothetical experiments explicitly to inspect interval coverage.
Same scientific action without the plot
Choose the observation set, estimator, and interval kind from the lists. Type the displacement count M and the inference temperature, viscosity, and radius. Read the table of estimated D and its interval, then N and its interval with the wording above, the inverse-bias note, and the identifiability family. The plots are a view of that same accepted snapshot.
Not modeled: localization error, blur, correlated or irregularly timed increments, and censoring (see BM-08 and kitchen mode); non-Gaussian increments; time-varying drift; polydispersity within one track set; wall effects; uncertainty in C without declared coverages; uncertainty in the gas constant itself.
Accepted calculation identity and limits
Run bm07-_R_6av5uiv5b_/run/1; snapshot 1; revisions {"input":1,"observer":0,"measurement":0,"estimator":0}. Host reference owner inference.bm07, scenario scenario-bm07-hidden-number.
Primary recording draws: 16385. Work for this request: 16385 draws, including 0 for hypothetical experiments. Retained primary recording: 65552 bytes. Reused worker recording: no.
Source identity: source:sha256:2b658a4b9461d836ec8ed1abf378bf46cab28ae0990ee15487493a19446add28. This host preview has no historical data importer or camera-noise fit. Browser arithmetic is not a claim of strict cross-engine WASM replay.
Each placement has separate settings and worker ownership. The same seed initially reproduces the same synthetic data; choose a different seed for an independent trial.
Open the inverse argument
The same positions, different questions
For independent, equally spaced Gaussian increments in d coordinates with known zero drift, average their squared lengths and divide by twice the number of coordinates and the time interval. The result estimates the diffusion coefficient.
When drift is fitted from the same observations, subtract the mean increment in each coordinate. The unbiased spread estimate uses M − 1 in place of M and q = d(M − 1). The centered maximum-likelihood estimate instead divides by M and must be rescaled before applying the same interval procedure.
An interval for a procedure
Under this ideal model, q times the unbiased diffusion estimate divided by the true diffusivity follows a chi-square distribution. Its quantiles provide the following conditional interval. More independent observations generally tighten uncertainty; changing the assumptions changes what the interval means.
At 95% coverage, the procedure covers the fixed true value in 95% of repetitions under its assumptions. That is not a 95% posterior probability for one parameter in one realized interval. The repeated-trial plots retain the misses as well as the successes.
Why an independent radius matters
With temperature and viscosity given, displacements identify the product Na. They cannot separately identify N and the radius a. Declaring a radius permits a conditional inversion, but deriving that radius from the same data using an assumed N would merely return the assumption.
Inversion also reverses the interval’s endpoints. It does not preserve unbiasedness: for the known-zero-drift estimate, the mean recovered-to-true number ratio is q/(q − 2) when q is greater than two. A point estimate can be finite even when that repeated-estimate mean does not exist.
Do not hide input uncertainty
The conditional interval treats the stated inputs as exact. The combined procedure uses the declared temperature, viscosity and radius intervals and adds their error probabilities with the diffusion interval’s error probability. This union-bound guarantee is conservative and does not require their statistical independence. It remains conditional on exact calibration, timing and the chosen gas constant in this preview.
Worked interval (readable without JavaScript)
| Quantity | Value |
|---|---|
| Diffusion interval | [0.331457, 0.578589] μm²/s |
| Modern-SI inversion (consistency check) | N̂ = 6.02213 × 10²³ mol⁻¹ |
| Inverse-bias factor at q = 100 | q/(q − 2) = 1.020408; an unbiased D̂ is not unbiased after inversion |
A synthetic recovery is not a new molecular count
This exercise draws a hidden number and generates data from it using a chosen gas constant. Recovering that number tests inference under the model; it does not measure a real suspension. With modern SI constants, the Avogadro constant is defined, so an observational inversion would be a consistency check rather than an independent determination. Reviewed historical observations and independent gas-constant measurements are not supplied here. This instrument uses ideal positions; the separate camera laboratory adds an explicitly later measurement model.
NIST: chi-square confidence limits for a normal variance · BIPM: SI units and defining constants
Next: keep the particle, change the camera, and test the inference →