The N-Back Working Memory Test, Explained
After reading this you will understand what the N-back task measures, how its hit-minus-false-alarm score is computed, and how to read your own numbers without fooling yourself.
What the task is, in one example
The N-back is the standard laboratory measure of working memory: your ability to hold recent information in mind and update it as new information arrives. Every 2.5 seconds a square lights up on a 3×3 grid and a letter appears. Your job is to compare the current stimulus with the one from N steps back and signal a match.
Consider a 2-back run. Suppose the positions form this stream, reading left to right:
top-left, center, top-left, right, center
At position 3 (top-left) you compare against position 1 (top-left). Match. Press A. At position 5 (center) you compare against position 3 (top-left). No match, so you stay quiet. That single act of comparing "now" against "two ago" while a new item is already arriving is the whole difficulty. You are not remembering a fixed list. You are running a moving window that shifts by one every 2.5 seconds.
The dual version runs two of these windows at once: one for positions (press A) and one for letters (press L). The two channels are scored separately, so a position match and a letter match can land on the same step or on different steps.
When to use it and when not
Reach for the N-back when you want a working-memory challenge that scales cleanly. Set N=1 and it is nearly trivial. Set N=2 in dual mode and most untrained adults score below 50%. That steep rise in difficulty makes the task useful for tracking your own practice over weeks.
Do not treat it as an IQ test. Whether N-back training raises general fluid intelligence is still argued about in the research literature. What nobody disputes is narrower: you get better at N-back itself with practice, partly through genuine memory gains and partly through strategy. Read your scores as scores on this task, not as a verdict on your mind.
The famous 2008 training studies used dual N-back specifically because it loads two channels at once and resists simple rehearsal tricks. Single-channel modes are still good practice, but the dual mode is the one the headlines were about.
The scoring formula and why it resists mashing
Each step is one of two kinds for a given channel: a target (it really matches N steps back) or a lure (it does not). Your response is either a press or silence. That gives four outcomes: a hit (press on a target), a miss (silence on a target), a false alarm (press on a lure), and a correct rejection (silence on a lure).
The score borrows from signal-detection theory. It rewards catching real matches and penalizes crying wolf:
Here hits is the count of correct presses on targets, targets is the total number of real matches in the run, false alarms is the count of presses on non-matches, and lures is the total number of non-matches. The first fraction is your hit rate; the second is your false-alarm rate. Subtracting one from the other is the "d-prime-flavored" logic mentioned on the tool page.
The subtraction is what makes mashing worthless. If you press every step, your hit rate is 1.0 but your false-alarm rate is also 1.0, so the score is (1.0 - 1.0) \times 100 = 0. Perfect silence also scores 0: hit rate 0 minus false-alarm rate 0. Only selective, accurate pressing moves the number up.
A worked example using the demo defaults
Scoring one position channel
Take a single 2-back run over a session with steps producing 6 real position targets and 14 non-matches (20 comparable steps). Suppose you press A correctly on 5 of the 6 targets and press wrongly on 2 of the 14 non-matches.
- Hit rate: 5 / 6 = 0.8333.
- False-alarm rate: 2 / 14 = 0.1429.
- Subtract: 0.8333 - 0.1429 = 0.6905.
- Scale by 100: 0.6905 \times 100 \approx 69.
Your position score is about 69. Now imagine you had grown impatient and pressed on 8 of the 14 non-matches while still catching 5 of 6 targets. The false-alarm rate jumps to 8 / 14 = 0.5714, and the score falls to (0.8333 - 0.5714) \times 100 \approx 26. Same hits, far worse score. The penalty for guessing is real and large.
In dual mode the letter channel is scored the same way over its own targets and lures, and the two channel scores are reported separately so you can see which one is dragging you down.
How the score behaves as you press more freely
The chart below fixes the hit rate at 0.8333 (5 of 6 targets caught) and sweeps the false-alarm rate from 0 to 1. Watch the score fall in a straight line, crossing zero when your false-alarm rate equals your hit rate.
The vertical marker sits at the crossover. Everything left of it is a positive score; everything right is negative. This is why a careful player who ignores uncertain steps often beats an eager player who presses on hunches.
Reading and interpreting your results
Anchor your expectations with a rough guide. These are typical ranges, not cutoffs, and they vary a lot between people and sessions.
| Mode | N | Typical untrained score | Reading |
|---|---|---|---|
| Single | 1 | 90 to 100 | Near ceiling; use only as a warm-up |
| Single | 2 | 60 to 85 | A real test of the moving window |
| Dual | 2 | 20 to 50 | The classic hard version |
| Dual | 3 | 0 to 30 | Very hard; expect many negatives early |
Expect a cliff between 1-back and 2-back. A 1-back run where you score 96 can drop to a 2-back run at 45 for the same person on the same day. That gap is the signal, not noise. It marks where holding one extra item exceeds your comfortable capacity.
Compare like with like. A dual 2-back score of 40 is not worse than a single 2-back score of 75; they are different tasks. The device keeps your best score per setting so you can track each combination on its own.
Common mistakes that cost you points
The biggest self-inflicted wound is pressing when unsure. Because lures outnumber targets (14 to 6 in the worked example), a single false alarm costs 1/14 \approx 0.0714 of score while a single hit is worth 1/6 \approx 0.1667. Two doubtful presses can wipe out a hard-won hit. When in doubt, stay silent.
A second mistake is verbal rehearsal in the letter channel. Repeating the last two letters aloud works for a while but collapses at 3-back and in dual mode, because you cannot rehearse letters and positions with the same inner voice. Better strategies bind each item to a mental image or a location, so the two channels do not compete.
A third mistake is chasing your best score every session. Working memory fluctuates with sleep, caffeine and stress. One low run does not mean you regressed. Average several runs at a fixed setting before you conclude anything.
Related tools on this site
The N-back sits in a family of short self-measures. If you want to isolate raw speed rather than memory, the Reaction Time Test reports your response latency in milliseconds. For a different working task under time pressure, the Mental Math Trainer drills arithmetic with per-operation accuracy, and the Doomsday Weekday Trainer exercises procedural memory for dates. To gauge sustained motor consistency instead of memory, try the Typing Speed Test or the timing jitter measured by the Tempo & Rhythm Tapping Test.
Frequently asked questions
Why did I score zero when I pressed every time?
Pressing on every step makes your hit rate and false-alarm rate both equal 1.0. The score subtracts one from the other, giving (1.0 - 1.0) \times 100 = 0. The formula is built so that guessing cannot earn points.
Can my score go negative?
Yes. If your false-alarm rate exceeds your hit rate, the difference is negative. At a hit rate of 0.8333 and a false-alarm rate of 1.0, the score is (0.8333 - 1.0) \times 100 \approx -17. A negative score means you pressed more indiscriminately than you caught real matches.
Is dual 2-back really that much harder than single 2-back?
For most people, yes. Single 2-back often lands at 60 to 85, while untrained dual 2-back commonly sits below 50, because you are running two independent moving windows and dividing attention between them.
Will practicing N-back make me smarter?
The evidence on transfer to general fluid intelligence is mixed and still debated. What is clear is that you improve at N-back itself with repetition. Treat gains as gains on this task unless a specific study says otherwise.
Does anything leave my device?
No. The task runs entirely in your browser, and your best score per setting is stored locally on the device you used.