Attention Training Technique (ATT) — Full 12 Minute Guided Session | Metacognitive Therapy

The Attention Training Technique happens entirely by ear, which makes the recording you use part of the exercise rather than a wrapper around it. Wells’ manual notes that recorded versions of sounds have been used in the implementation of ATT[1] — so audio delivery is documented, not a workaround. What the manual does not do is bless any particular recording, which leaves the question of what a faithful one has to contain.

The video above is a full twelve-minute session, so you can hear the shape before reading about it.

The checklist

A recording that delivers the documented exercise has to have all of these. Most audio labelled “attention training” has some of them.

RequirementWhat the protocol specifies
Several sounds at onceBetween six and nine, depending on the level of demand required[1]
Sounds at different distancesAt least three competing sounds in the room, two more nearby, two in the far distance[1]
A voice that names themThe script directs attention to a specific sound, with an instruction to refocus when attention is captured[1]
Three phases, 5–5–2About five minutes selective, five switching, two divided — around twelve minutes total[1]
An accelerating switchRoughly every ten seconds, increasing to one sound every five[1]
A different soundscape next timeSounds and arrangement varied between sessions, to offset the effects of practice on task difficulty[1]

What disqualifies a recording

Ambient tracks with no voice. Rain, forest, café noise — pleasant, and not the exercise. Without a voice naming targets there is nothing directing attention, and directing attention is the entire task[1].

Anything sold as relaxation. The manual is explicit that ATT should not be employed as a distraction, avoidance, or symptom-management strategy, and is not intended as an emotion-management strategy[1]. A recording whose stated purpose is to calm you down is aimed somewhere else.

Guided meditations that mention sounds. Attending openly to whatever sounds arise is a different instruction from being told which sound to hold and when to leave it.

A single track you replay. The protocol asks for variation between sessions[1], precisely so the exercise does not become easier. One memorised recording works against its own design.

Binaural beats and similar. Nothing in the protocol involves tones intended to influence brain rhythms; the exercise is a task you perform, not a signal applied to you.

Why good ATT audio is hard to find

The technique comes from a clinical manual, where the sounds were whatever the consulting room offered and the script was read live by a therapist[1]. Nothing in that origin produces a distributable recording — the manual’s own answer to the question is a single closing line pointing readers to the MCT Institute’s website for a recorded version of the ATT[1]. Most publicly available audio was made by individual practitioners or channels working from the published description, and quality varies in ways a listener cannot easily audit — the specification above is what to check against. The ones we have found and checked, in eight languages, are listed with direct links on ATT on YouTube.

What “distance” means in a recording

The near/far structure is the part recordings most often flatten. In a room, the sounds are genuinely at different distances; in a stereo file, that has to be reconstructed — otherwise seven sounds arrive at one apparent location and the layout the protocol specifies is lost.

The manual’s own layout has three distance bands and includes potential sounds — locations you are directed to attend to where no sound may be present at the time[1]. A recording that only ever names things you can currently hear is dropping one of the specified conditions.

What to do with the recording once you have one

The audio is the delivery, not the practice. The protocol’s instructions still apply: eyes open on a fixation point[1]; thoughts treated as additional noise and not resisted[1]; practice scheduled when you are not in a state of anxiety or acute worry[1]; and no expectation that it becomes comfortable, since it is built to stay demanding[1].

For the sounds themselves, see what sounds do you use for ATT?; for what the voice is doing, see what the ATT script says.

For where to get one free, see free ATT sessions and ATT on YouTube.

Questions and answers

Is there a free ATT audio recording?
Recorded delivery is documented in Wells' manual, and full sessions exist publicly — including the twelve-minute session on this page. What is scarce is recordings that state which protocol they follow.
What makes an ATT recording faithful to the protocol?
Several simultaneous sounds at different distances, a voice that names them and directs attention between them, the 5–5–2 phase structure, and a switching phase that accelerates from about ten seconds to about five.
Can I just use nature sounds or focus music?
No. Those give attention nothing to be steered between and no voice directing it. Without the naming and switching instructions it is not the exercise, whatever the sounds are.

References

  1. Wells, A. (2009). Metacognitive Therapy for Anxiety and Depression. Guilford Press. Chapter 4: Attention Training Techniques.

Published by the makers of Heed, an app for practicing ATT-style sessions. iPhone · Android. Heed is a self-practice tool inspired by ATT research — not therapy, and not a medical device.