Psychophysics Bench Developer Reference
Maintain the browser psychophysics runtime: UI contract, configuration, state, staircases, the Bayesian tracker, the accelerated ladder, stimulus rendering, response handling, charts, exports, URL state and Library integration.
Overview
The bench is a static page plus three plain scripts: a numeric core that touches no document, a register of what every figure rests on, and a page half that runs the trial loop and draws. Two things it reports were on the wrong axis and are the reason for the rebuild — contrast was built in 8-bit code values with no transfer curve anywhere, and spatial frequency was cycles across the patch while the slider said cycles per degree.
1. Code map
+| File | Role |
|---|---|
tool/general/perception-and-neuroscience/psychophysical-experiment-engine/index.html |
SEO shell, masthead, provenance strip, tab rail and the six panels: Lab, Method, Paradigms, Data, Export, Reference. |
js/tool/pe/pe-core.js |
Every calculation and nothing that touches a document: the
transfer curve, contrast in luminance, viewing geometry, the
stimulus renderers, the staircases, the psychometric fit, the
detection metrics and the bootstrap. Exposes
window.AAPECore, and the test suite runs it
directly. |
js/tool/pe/pe-sources.js |
The dataset register: thirteen entries with a status of
verbatim, computed, synthesised or absent. Exposes
window.AAPESources. |
js/tool/psychophysical-experiment.js |
The half that has a document: state, the timed trial loop, the
readouts, the staircase trace and the exports. Exposes
window.AAPEEngine and
window.AAPsychophysicalExperiment. |
js/tool/psychophysical-experiment-views.js |
Method, Paradigms, Data, Export and Reference. Holds no numbers of its own except the historical contrast path it reconstructs so the page can show what changed rather than assert it. |
css/anz-shell.css,
css/psychophysics-shell.css |
The shared workbench slice and this route's sheet. The route
loads the slice instead of analyzer.css. |
tests/colorimetry/pe-harness.js,
pe-reference.test.js, pe-sources.test.js |
60 tests that load the shipped core and register under a stubbed global, so every assertion runs the code the page runs. |
2. The two axes
+Contrast. Every contrast is Michelson contrast in relative luminance
about a stated mean: the peak and trough are placed as
L × (1 ± m) and encoded back
through the sRGB transfer curve as the last step. Placing the swing that way
returns exactly the contrast asked for; doing the same arithmetic on code
values does not.
The page built the grating as
0.5 + 0.5 × c × cos(…)
scaled by 255, with no transfer curve anywhere — the word
gamma appeared twice in the whole engine, both times as the guess
rate of the Bayesian staircase. Swept through the sRGB curve, what that
delivered was:
| Asked | Codes sent | Delivered | Error |
|---|---|---|---|
| 0.02 | 130 / 124 | 0.0510 | +155% |
| 0.05 | 133 / 121 | 0.1018 | +104% |
| 0.10 | 140 / 114 | 0.2183 | +118% |
| 0.50 | 191 / 63 | 0.8258 | +65% |
| 1.00 | 255 / 0 | 1.0000 | 0% |
A Gabor detection threshold sits between about 0.005 and 0.02, which is where the error was worst. It is not a scale factor that can be divided out afterwards, because it depends on the mean the excursion was taken about — and the old export did not record the mean, so sessions on the old axis cannot be converted even in principle. They are marked instead.
Frequency. degreesPerPixel() is
atan(pitch ÷ distance), and every frequency on
the page derives from it. The slider was labelled c/deg, the documentation
printed the conversion, and the shader-side expression was
cos(2π · sf · x ÷ patchWidth)
— cycles across the patch. Searching the old engine for a viewing
distance, a pixel pitch or a degree returned nothing.
The Nyquist limit of the pixel grid is half a cycle per pixel expressed in c/deg, and the Lab warns above it: past that the display cannot carry the frequency being asked for, so what reaches the eye is an alias at a different frequency and a different contrast.
3. UI contract
+Ids use the pe- prefix. The Lab rail carries the task, the
stimulus, the viewing geometry, the staircase and the timing; the stage,
the response keys, the staircase trace and eight readouts sit beside it.
#pe-paradigm— two options, because three of the old five could not be scored (§7).#pe-cpd,#pe-orientation,#pe-sigma,#pe-mean-code— the stimulus. The envelope sigma is in degrees and becomes pixels through the geometry, so it means the same thing at any viewing distance.#pe-distance,#pe-pitch— the two numbers without which a spatial frequency is not a spatial frequency.#pe-stair,#pe-step-db,#pe-start,#pe-reversals,#pe-trials— the procedure. The step is in decibels of contrast because it is a ratio.- The response keys are ← and →, and the on-screen buttons are 44 px tall so a participant can use them on a phone. Esc stops a run.
The Contrast slider is gone. It was read into the configuration, written into
every saved preset, put into every shared link and set by all six built-in
presets — and cfg.contrast was never read by anything
that rendered a stimulus or ran a trial. The contrast that matters comes
from the staircase, and its starting point has its own control.
4. Configuration and state
+readState() reads fifteen fields off the rail into a flat
object. Nothing else reads the DOM for a value, so a test can drive the core
with a plain object and get the page's arithmetic.
The run's own state is separate: the ladder or tracker, the trial rows, the current contrast and what was true on the current trial. It is rebuilt from scratch on every start, so a second run cannot inherit a posterior or a reversal list from the first — which is the sort of thing that makes a second session quietly better than the first.
5. Adaptive procedures
+The step is multiplicative. Staircase moves by a constant
ratio given in decibels, so the resolution is the same fraction of the
contrast at every level. The page stepped by a fixed 0.05 on a range of 0.01
to 1: a step larger than the threshold cannot resolve the threshold, and
near 0.01 the ladder jumped between 0.01 and 0.06 — a factor of six in
one trial.
The reversal average takes an even count.
thresholdFromReversals() discards the warm-up and then drops
the oldest of what remains if the count is odd, because up-runs and
down-runs are not symmetric about the threshold (Levitt 1971). The mean is
geometric, because the ladder steps by ratios and an arithmetic mean of
ratios biases upward. The readout says how many were used and how many were
discarded.
Two procedures ship under their own names. The Bayesian tracker is on a log grid with a Gaussian prior of stated width; the page called it QUEST, and Watson & Pelli's method is defined on log intensity with an informative prior, which a linear grid under a flat prior is not. The accelerated ladder halves on a reversal and doubles after three in a direction; the page called it PEST, and Taylor & Creelman's is a sequential Wald test. Both are useful; neither is what it was called, and the Data register carries both as stand-ins with what they would take.
A test runs each transformed up-down rule for 600 trials against an exactly-specified observer, thirty times, and measures the proportion correct over the settled half: 1-up/2-down lands at 70.5% against a target of 70.7%, and 3-down/1-up at 79.4% against 79.4%.
6. Stimulus runtime
+gaborPatch() and gratingPatch() return the pixels
and what the display could actually deliver. The envelope multiplies the
contrast, so the modulation at the centre is the contrast asked for, and the
sigma is a real stimulus parameter rather than a cosmetic one — it was
hard-coded to a quarter of the patch width.
deliverContrast() caps the excursion at what fits about the
chosen mean: the trough cannot go below zero and the peak cannot exceed
white. The Lab reports delivered / asked on every run and raises a
strip when they differ, because a threshold measured against a clipped
stimulus is a threshold for something else.
measureContrast() reads the Michelson contrast back out of the
rendered pixels through the same transfer curve. The renderers are held to
that measurement in the tests rather than to their own arithmetic, which is
the check the old code could never have passed.
7. Trial loop and scoring
+scoreTrial(paradigmId, trial) takes the trial rather than
reaching for module state, and returns true,
false or null. The third value is the
point: a trial that cannot be scored has to be distinguishable from one
answered wrongly.
The page set the stimulus interval to −1 for every task
that was not two-interval and then compared a response of 0 or 1 against it,
which is false. So m-AFC, same/different and odd-one-out marked
every trial wrong and drove the staircase to its ceiling. Driven in a
browser, four sessions of 24 trials each:
| Task | Accuracy | Where the ladder ended |
|---|---|---|
| 2AFC | 0.42 – 0.46, varies | 0.90 – 1.00 |
| Yes / no | 0.46 – 0.67, varies | 0.58 – 0.93 |
| m-AFC | 0.000 every run | 1.00 every run |
| Same / different | 0.000 every run | 1.00 every run |
| Odd one out | 0.000 every run | 1.00 every run |
Those three are withdrawn, each with what it would take to bring it back. If
scoreTrial ever returns null during a run the loop
stops rather than record a trial nothing can interpret.
8. Fit, thresholds and detection metrics
+fitPsychometric() fits
ψ(x) = γ + (1 − γ − λ) F(log x)
with F logistic. γ comes from the task, not from
the fit; λ is fitted over a small range, because an unmodelled
lapse biases both threshold and slope (Wichmann & Hill 2001, part I
— which the page cited while doing exactly that). The search is a
coarse grid followed by coordinate descent, not a 31 × 21
grid whose slope pegged at an integer boundary.
thresholdAt(fit, p) returns the contrast at a stated
proportion and NaN at or below the guess rate. The page's fit
ran from 0 to 1, so its threshold was the contrast at 50% correct —
which in a two-interval task is chance, the contrast at which the observer
is guessing, while the staircase was converging on 70.7%. Two different
quantities were printed side by side as one.
calcSDT() returns d′, criterion and β for yes/no
only, and a reason otherwise: a two-interval task has no stimulus-absent
trials, so it has no false-alarm rate and no d′ of this form. The
extreme-proportion correction is the 1/2N rule (Macmillan & Kaplan
1985); the page credited it to Hautus 1995, which is the paper comparing it
unfavourably with the log-linear correction. The inverse normal is
Abramowitz & Stegun 26.2.23 with its input clamped to [0.001, 0.999],
which caps |d′| near 6.18 — a reading that reaches the
ceiling says so.
bootstrapThreshold() resamples trials, refits and takes the
2.5 and 97.5 percentiles. It is labelled a trial bootstrap because an
adaptive staircase chooses each trial's contrast from the responses before
it, so resampling trials as independent narrows the interval.
9. Exports and URL state
+The CSV carries a header before the rows: the contrast unit, the mean
luminance and the code it came from, the frequency unit, the viewing
distance and pixel pitch and the degrees per pixel they give, a line saying
the transfer curve is the sRGB standard rather than a measurement of this
screen, the paradigm, the staircase with its target, and
reportable.
Every line is there because its absence made an old file uninterpretable. A
Michelson contrast cannot be reconstructed without the mean it swung about,
and the old export recorded neither the mean nor the unit — which is
why old sessions are marked rather than converted. Two columns ship always:
contrast is what was asked for and delivered is
what the display could reach.
The share link carries the whole rail as a query string, and
loadURL() reads it back into the rail when the page loads.
10. Public API and library binding
+window.AAPECore— the numeric surface. No DOM, so it can be called from a test, a worker or another tool.window.AAPESources— the register:all(),byId(),byStatus(),reportable(),counts().window.AAPEEngine—getState(),getRun(),threshold(),targetProportion(),csv(),sessionObject(),refresh().window.AAPsychophysicalExperiment— thegetState/restoreStatepair the Library capture already uses. Kept under its old name so nothing downstream breaks.
11. Tests
+npm run test:colorimetry runs 60 tests for this tool, in
tests/colorimetry/pe-reference.test.js and
pe-sources.test.js. They load the shipped
js/tool/pe/pe-core.js and pe-sources.js under a
stubbed global, so every assertion runs the code the page runs.
Each group was proved by injecting the original defect back into the core and confirming the group fails. Twenty mutants, twenty caught: the grating back in code values; the frequency back in cycles per patch; degrees per pixel dropping the viewing distance; an unscoreable trial scored as wrong; the withdrawn tasks back in the menu; a stand-in dropping what it would take; the step additive again; the reversal average taking an odd count and taking an arithmetic mean; 1-up/2-down stepping down on one correct trial; the Bayesian tracker on a linear grid and under a flat prior; the fit losing its guess rate and its lapse rate; a threshold reported at chance; the extreme-proportion correction dropped; d′ computed without noise trials; the criterion losing its sign; the contrast cap unreported; and the bootstrap dropping its caveat.
Two of those mutants survived the first pass, and both were the test's fault rather than the code's: the criterion check used a symmetric operating point where a sign flip is invisible, and nothing pinned the fit's guess rate to the task. Both now have tests that discriminate.
By hand, in a browser: run both paradigms and confirm the proportion correct sits near the ladder's target; move the viewing distance and confirm the Nyquist figure and the alias warning follow; raise the background until the contrast caps and confirm the strip appears and delivered/asked drops below one; and check the page at 390 px, in both themes and in Hindi.