Blog

Virtual birds were easier to find than real ones

· News, Projects

We spent two days at the Great Exhibition Road Festival running a listening game we built for SONICOM, the EU research project on immersive audio we are part of. Visitors put on a Quest 3, picked up a snowglobe with a bird inside it, and walked to the spot in the room the bird was calling from.

Every time somebody picked up a globe, the software flipped a coin. One outcome played the call through a loudspeaker standing at that spot in the room. The other rendered it inside the headset, positioned at the same point in space. Nobody was told which they were hearing.

We went in expecting to measure how far short of a real loudspeaker the headset falls. We got the opposite result: people found the virtual birds more reliably than the real ones, and took the same amount of time doing it.

The Bird Song Chamber

We ran it in a teaching room at Imperial College London, draped in black cloth, with the snowglobes waiting on a white table in the middle. The room held eight physical loudspeakers: four marking the dropzones where the birds live, and four more for forest ambience. Every one of them had a virtual twin, rendered through the Meta XR Audio spatialiser and played over the Quest 3’s built-in speakers, with passthrough on so people could see the real room the whole time.

Four snowglobes sat on a table, one bird each. Picking one up started its call. The task was to carry it to the dropzone the call was coming from and set it down. Four birds placed correctly finished a stage, and there were three stages:

  1. Tutorial. Narration walks through the task and the correct dropzone glows green when a globe is picked up, so vision carries the answer.
  2. Direct sounds only. The visual hint is gone. Bird calls are the only sound in the room.
  3. Masking sounds added. The hint stays off and the forest ambience starts, competing with the calls.

Between stages the four birds swap positions, so nobody could carry a memorised map forward.

A Max/MSP patch ran the game and logged everything: head position and rotation at about 20 Hz, the position of every snowglobe, each pick-up and put-down, and the time from grabbing a globe to placing it correctly. Two days of that came to 198,563 events across 77 sessions, 71 of which were complete enough to analyse.

The result

Placement accuracy split cleanly by condition once the visual hint was switched off.

Real loudspeakerVirtual, rendered by the Quest 30%25%50%75%100%85.6%n=16085.6%n=195Stage 0tutorial, visual hint on83.4%n=15194.5%n=181Stage 1direct sounds only78.3%n=16193.3%n=179Stage 2masking sounds addedguessing would score 25%

Across the two scored stages, virtual audio was right 93.9% of the time against 80.8% for the real loudspeakers. That is 13.1 percentage points, with a bootstrapped 95% confidence interval of 5.7 to 20.7 points, and odds of a correct placement 3.7 times higher.

The tutorial stage is the control we did not plan for. While the correct dropzone was glowing green, both conditions scored exactly 85.6%. Vision was doing the work and the audio made no difference at all. The gap only appears once we take the hint away, which is our best evidence that the ears are what separated the two.

23 peopledid better with virtual33 peoplescored the same6 peopledid better with real62 participants with trials in both conditions

Speed came out flat. Median search time was 4.8 seconds with a real speaker and 5.3 with a virtual one, and the per-participant paired test shows nothing (p = 0.59). People were not hesitating with the virtual audio and then getting lucky. They walked straight to it.

The misses tell the same story from the other side. 98% of failed real-audio placements ended up within arm’s reach of one of the other three dropzones, a median 2.08 m from the right one. A real-audio miss also cost 2.09 m of walking and 406 degrees of head-turning, well above a successful trial. Those were people searching hard and still landing on the wrong bird’s home.

Our reading of it

The explanation we find most convincing is that the virtual bird sits exactly where the dropzone is. Both are generated by the headset, in the same tracking frame, so sound and sight agree perfectly and the choice becomes easy. A real loudspeaker sits at a physical spot that is only approximately where the AR dropzone appears, and headset drift over a session pulls the two further apart. That predicts what we see: no effect while the visual hint is present, a gap once it is removed, and a wider gap for distant targets, where a few degrees of misalignment turn into a metre of error.

The masking stage points the same way. Real accuracy drops to 78.3% once the ambience comes on while virtual holds at 93.3%. A real bird call and real ambience share the room’s reverberation and the same loudspeaker character, so they compete directly. A headset-rendered call stays separable and comes through the mix.

Loudspeaker directivity is probably in there too. The Genelecs are directional and sat on cloth-covered tables in a hard room, so anybody off-axis heard a reflection-dominated version of the call. The binaural renderer has no such problem, because its cue set is idealised.

The baseline this sets

SONICOM’s founding ambition is to blend the real and the virtual so that it is not possible, from an auditory point of view, to tell them apart. This is not that test, because we never asked anyone to judge which was which. What it does give us is a floor, and the floor is higher than we assumed.

The virtual condition here was not a laboratory best case. It was the stock Meta XR Audio spatialiser, a generic head-related transfer function, no room matching, no personalisation, played over the open-ear speakers on a consumer headset. That configuration was already good enough to beat a physical loudspeaker at a task that only rewards getting the direction right. Anecdotally, very few visitors realised the sounds were sometimes virtual at all.

For scenario design, that means we can stop treating spatial placement as the thing limiting how convincing a mixed-reality scene feels, at least for rooms this size and layouts this sparse, and put the effort into what SONICOM is actually researching: personalised HRTFs, reverberation matching and source directivity. Building those into the Bird Song Chamber is what comes next, by swapping the stock spatialiser for the consortium’s Binaural Rendering Toolbox through the Unity wrapper we already wrote for it.

The Bird Song Chamber was built by Yuli Levtov and Ragnar Hrafnkelsson at Reactify with Becky Stewart at Imperial College London.

eu-logo

This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 101017743.