Skip to main content
Auric Artisan · Documentation

Psychophysics Bench User Guide

Use Auric Artisan Psychophysics Bench to build repeatable browser-based threshold experiments with adaptive staircases, a Bayesian tracker, an accelerated ladder, psychometric fits, SDT metrics and local exports.

Published: May 25, 2026 Updated: May 25, 2026 Category: Guide Author: Chirag Bansal
Back to Documentation Auric Artisan Home

Overview

The bench measures a contrast threshold: the faintest grating you can reliably tell apart from a blank. You set the stimulus and how far you are sitting from the screen, answer a few dozen trials with two keys, and it reports the contrast you were right at a stated percentage of the time — in luminance, which is the unit a threshold has to be in to mean anything.

Table of contents

  1. 1. Open the bench
  2. 2. The two tasks
  3. 3. Tell it where you are sitting
  4. 4. The stimulus
  5. 5. How the contrast is chosen
  6. 6. Running a session
  7. 7. Reading the result
  8. 8. Taking the data with you
  9. 9. What this is not

1. Open the bench

+

Go to Psychophysics bench. Everything runs on your own machine; nothing is sent anywhere.

The Lab tab is where you run. Method shows what a contrast is on this page and what each staircase is estimating. Paradigms lists the tasks and says which ones can be scored. Data is the register of everything a threshold rests on. Export and Reference are what you take away and what it is built on.

2. The two tasks

+

Two-interval forced choice. Two intervals go by, marked on screen. The grating is in exactly one of them. You press ← for the first or → for the second. If you cannot see it, guess — guessing is part of the method, and chance here is 50%.

Yes / no detection. One interval, and the grating is there on about half the trials. ← for yes, → for no. This is the task that gives you d′, because it is the only one with trials where nothing was shown.

The menu used to offer five. Three of them — m-alternative, same/different and odd-one-out — scored every trial as wrong, whatever you pressed, and drove the contrast to maximum. They are withdrawn, and the Paradigms tab says what each would need to come back.

3. Tell it where you are sitting

+

This is the step people skip, and it is the one that makes the numbers mean something. Spatial frequency is in cycles per degree of visual angle — how many light-dark cycles fall into one degree of your field of view — and that depends on how far away you are and how big your pixels are.

Viewing distance is from your eyes to the screen. 57 cm is the conventional default because at that distance one degree is almost exactly one centimetre, which makes everything else easy to check.

Pixel pitch is the physical size of one pixel. Divide your screen's width in millimetres by its width in pixels: a 24-inch 1920×1080 monitor is about 0.277 mm, a typical laptop about 0.19 mm. The default, 0.248 mm, is a 1080p 21.5-inch panel.

Get these two roughly right and your threshold is comparable with published work. Leave them wrong and the frequency you think you measured is not the frequency you measured.

4. The stimulus

+

A Gabor is a grating fading out under a soft Gaussian blob — the standard stimulus for contrast detection, because it is narrow in both position and frequency. A grating in an aperture is the same modulation with a hard circular edge.

Spatial frequency chooses how fine the stripes are. Human contrast sensitivity peaks somewhere around 2–5 c/deg, so that is where your threshold will be lowest. Orientation rotates them. Envelope σ is how big the blob is, in degrees — so it stays the same size in your visual field when you move.

Background sets the grey the contrast swings about. It matters more than it looks: contrast is a swing either side of that grey in luminance, and a bright background leaves less room to swing before the display runs out. When the contrast being asked for will not fit, the bench caps it and tells you — the Delivered / asked readout drops below 1.00 and a strip appears. A threshold measured against a clipped stimulus is a threshold for something else.

5. How the contrast is chosen

+

You do not pick the contrast — the staircase does, moving it up when you get one wrong and down when you get a run right, so it spends its trials near your threshold rather than wasting them where you can obviously see or obviously cannot.

1-up / 2-down is the usual choice: two correct in a row lowers the contrast, one wrong raises it, and it settles where you are right 70.7% of the time. 3-down / 1-up settles at 79.4% and takes longer. Those two percentages are the answer — a threshold is a contrast at a stated percentage correct, and the readout names it.

The other three are honest about what they are. The Bayesian tracker and the accelerated ladder are this tool's own procedures and settle at no stated percentage. Constant stimuli does not adapt at all.

Step is in decibels because it is a ratio, not a fixed amount: 2 dB is a factor of 1.26 at any contrast. That is what lets the ladder resolve a threshold of 0.01 as precisely as one of 0.5.

6. Running a session

+

Press Practice first — a dozen trials to learn the rhythm, not recorded. Then Run.

  • Look at the centre of the stage and keep looking there.
  • A cross appears, then the interval or intervals, then it waits for you.
  • Answer with ← and →, or the two buttons.
  • Guess when you are unsure. Do not wait to be certain — the method assumes you are guessing near threshold, and hesitating biases the result upward.
  • Esc stops.

The run ends at whichever comes first: the reversal count you set, or the trial ceiling. The Bayesian tracker and constant stimuli have no reversals, so they run to the ceiling. Sixty to a hundred trials is a normal session. Dim the room, keep the screen brightness fixed, and do not change the viewing distance halfway through.

7. Reading the result

+

Threshold, Michelson is the answer: the contrast at which you were right the percentage of the time shown next to it. A healthy young observer at 4 c/deg in a dim room lands somewhere around 0.005–0.02 on a calibrated display; a bright room, an uncalibrated screen or a glance away will all raise it.

Reversals used says how many turning points went into the average and how many there were. The first two are warm-up and are discarded, and the count used is always even — the ladder's up-runs and down-runs are not symmetric, so an odd number would lean one way.

The trace shows the contrast on every trial, on a log axis. Grey dots are the discarded warm-up reversals, red dots are the ones used, and the gold dashed line is the threshold. A good run wanders up and down in a narrow band near the end. One that slides to the top and stays there means you could not see the stimulus at any contrast — check the background, the frequency and the alias warning.

8. Taking the data with you

+

The CSV has one row per trial and a header that says what the numbers are: the contrast unit, the background luminance the swing was about, the frequency unit, your viewing distance and pixel pitch, and a line saying the transfer curve is the sRGB standard rather than a measurement of your screen.

That header is the point. A contrast without the grey it swung about cannot be reconstructed later, and a frequency without a viewing distance is not a frequency. A file outlives the page that made it.

The JSON carries the same thing plus the settings and the threshold. The link button copies your whole setup into a URL, so you can send someone the exact configuration.

9. What this is not

+

Your display is not calibrated. The bench assumes the sRGB transfer curve, which is a standard, not a measurement of the panel in front of you. A real laboratory measures its own screen with a photometer. Until you drop a measured table into the Data rail, treat your thresholds as comparable with each other rather than with published values.

It is not a clinical instrument. Nothing here diagnoses anything. If you are worried about your vision, see an optometrist.

Some of it is this tool's own. The Data register marks every method as verbatim, computed, a stand-in or absent, and says what a stand-in would need to become the real thing. Two things the page used to advertise — a Weibull fit and an ROC curve — were never implemented at all, and are listed as absent rather than quietly dropped.

Old sessions are on a different axis. Before this rebuild the contrast the page reported was a code value rather than a quantity in light: a requested 0.02 reached the eye as 0.051. Those files cannot be converted, because the conversion depends on a background the old export did not record. They are marked rather than rescaled.