Every visual cue fails the moment the thing you need is behind you. Putting the remote expert's voice in 3D, and adding audio beacons, solved a problem the visual work could not.

The two systems earlier in this series shared visual cues: a reconstructed room, then gaze rays and hand meshes. Both work well, and both have the same blind spot, literally. A visual cue can only help you if it is already in your field of view. If the object you need is behind you, a hand pointing at it is a hand you cannot see.

Sound does not have that problem. Human hearing is 360 degrees by default.

This study was led by Jing Yang, with Huidong Bai, Amit Barde, Gábor Sörös and Mark Billinghurst, and I worked on it as second author. It ran alongside the gaze and gesture work and asked the complementary question: what happens if the expert can point with sound?

The setup

The remote expert works inside a 3D virtual replica of the local space and teleports around it rather than walking, which means they can be standing next to something the local worker has not reached yet. The local worker wears an AR headset, and their egocentric view is always live-shared back to the expert.

The remote expert teleports through a virtual replica of the local space, while the local worker's view streams back continuously.

Two things get spatialised. The expert's voice is positioned in 3D, so it arrives from where they are standing in the shared space rather than flatly in both ears. And the expert can drop auditory beacons, sounds attached to a location, which the worker perceives as coming from that spot in the room.

On top of that sit the visual cues from the earlier work: the head frustum showing where the expert is looking, and shared hand gestures.

The head frustum and the hand gesture, as shared from the remote side to the local side.

The task was search: find small objects that were deliberately occluded, in a large office. Search is the right task for this question, because it is precisely the case where the answer is usually not where you are already looking.

Findings

Spatialised audio beat non-spatialised audio on spatial perception, significantly, for finding small occluded objects. Being able to tell roughly where a voice is coming from is enough to orient you, before anyone says anything useful.

But audio alone gives you direction, not precision. The spatial cues told workers the general layout and roughly where to search. They did not pin down the object. That last step still needed vision.

The head frustum was the standout visual cue. Not the hand gesture, the frustum: the wireframe cone showing where the expert's head is pointed. It communicates viewing direction and field of view at a glance, and combining it with the spatial audio significantly improved task performance, social presence and spatial perception together.

The line from the paper that stuck with me is that some participants said the combination made them feel like they were interacting with a real person. Two studies in this series now produce that same sentence from different directions, which suggests it is measuring something real rather than novelty.

Why it belongs in the arc

The result is a division of labour that is cleaner than I expected. Audio is good at coarse and omnidirectional: it turns you toward the right part of the room from any starting orientation. Vision is good at fine and directional: once you are facing the right way, it puts your hand on the object.

That is not two redundant channels. It is two halves of a single act of pointing, and the earlier visual-only systems were doing the second half without the first.

It also completes the sensory story for this series. Part 1 built the channel, part 2 worked out which visual cues carry intent, and this one adds the dimension vision cannot cover. The remaining question is whether any of it actually teaches anybody anything, which is where the series goes next.

The paper is The Effects of Spatial Auditory and Visual Cues on Mixed Reality Remote Collaboration, Journal of Multimodal User Interfaces 2020, by Jing Yang, Prasanth Sasikumar, Huidong Bai, Amit Barde, Gábor Sörös and Mark Billinghurst. The PDF is here and there is a video preview here.

Read the next post in this series here:

You might also like …