IrSOb – The Iridescent Sound Object

A model and architecture for 3D sound object creation and binaural reproduction in VR spaces.

Article by Philippe Druez
Share this article
go to article
Figure 1 — Experiencing the IrSOb Heidelberg Press Sound Project at Orpheus Institute.


On The Poetics of Space .…

‘The poetic image is a sudden salience on the surface of the psyche’ and
has to be a quest based not on causality, but in ‘reverberation’ and ‘in
this reverberation, the poetic image will have a sonority of being’.

Gaston Bachelar


Abstract

IrSOb (Iridescent Sound Object) proposes a theoretical, object-oriented model and modus operandi for capturing real-world 3D sounds and constructing synthetic 3D virtual-reality sound objects. A proof-of-concept implementation enables participants to experience these sound objects either within a wavefield spatial-audio laboratory or at any location using a dedicated hardware system for binaural reproduction over headphones.

Due to its distinctive characteristics, this new method offers added value compared to current techniques in 3D gaming, sound art, performance, theatre, and scientific research involving embodiment. It also provides a valuable tool for the realistic recreation and archival preservation of musical instruments and a wide range of sounding objects. Thanks to its simplicity and coherent workflow, the method opens new creative possibilities for sound artists and composers working with experimental music.

The Iridescent Sound Object is part of my research at the eXtended Reality and Human Interaction Labs at IPEM (Institute for Psychoacoustics and Electronic Music), within the Systemic Musicology Section of the Department of Art, Music and Theatre Studies at Ghent University (UGent), and at KASK School of Arts and Conservatory Ghent.

Introduction

The use of spatial sound in the gaming and the movie industry is already well established. Major entertainment players made important progress in delivering spatial audio for cinema or home theatres and into the headphones of the consumer or even the museum visitor. The result delivered by these immersive experience systems can be quite impressive and are available from a wide range of manufacturers1.

This specific part of my PhD resulted in a method for acquiring faithful 3D sound captures or create non-existing VR sound imagery; a method for recording the 3D sound object; a Max for Live device for reproducing the 3D sound object in a 3D lab environment using Qualisys MoCap2 (XRHIL in this case) and Earis, a for the purpose developed and built hardware system using binaural headphone technology that can be used anywhere.

This article covers the theoretical concept and method of IrSOb and its relation to Ambisonics, an overview of related 3D recording technologies, the architecture of the Earis hardware and its intrinsic iRose spatial calibration system and in brief, The Heidelberg Press Sound Project (discussed further) as use case as demo at Orpheus (Figure 1).

IrSOb, the iridescent sound object's methodology is not to be confused with the object-based concept where audio designers place a sound as a "point" in 3D space (instead of channel-based by assigning a sound to a specific speaker)3, The Iridescent Sound Object is a VR 3D sound object in itself that has a “volume with tangible boundaries and can be placed in a VR 3D sound space.

Unlike omnidirectional point sounds (some treated with equalising for better orientation perception), IrSOb offers you how you hear a sounding object in real life as VR experience. These sound objects sound different when listened at from different angles. Additionally it will reflect different sounds to its “environment” and thus create a much more accurate and lively reflecting and reverb ambient experience (future extension). The development started at the XRHIL wavefield lab of UGent (IPEM) but is technology independent. As a proof of concept I developed a basic device to experience the IrSOb everywhere.

Figure 2 — L’image de Mon Père. The Sound Frame producing a 2D sound image. Insert: The four-mic recording setup.
Figure 3 — AcouStasis concept.

How did I get here…

When visiting the Immersive Lab at AP Hogeschool Antwerpen I saw a volumetric capture device of the photogrammetry studio and was thinking of replacing the cameras with microphones to capture a "real" 3D sound image of an object inside it. This could be a person singing, a violin playing, robins in a cage or a Bernina sewing machine, but nothing larger. Besides the complex and not so portable structure, how to define the number and placement of (and what kind of) microphones in what could be called a “sonogrammetry” device and make it scalable. Capturing a real 3D sound image of a violin performance looked feasible to me but that of ants at work in their nest or a Formula One Ferrari seemed as significant to me and needed a different approach.

I have the unfortunate tendency of transmogrifying any tiny idea into octopus-monsters with dozens of spin-off tentacles and deep holistic and futuristic visions. But the lack of a large studio, a team working for me and the requisite funding reduced this project to a workable field resulting in what became the IrSOb project. That said, the project itself is a bit of a spinoff of my initial PhD research and even more, the hardware developed is a side dish as proof of concept of this spin-off. So began my groundwork on how to capture any object outside a lab environment (portable) and with a minimum of construction hassle. This resulted in the iridescent sound object model on which I will throw light in this paper.

Before I started my research on 3D sound in 2019, I created a sound-art work with the title l'Image de Mon Père (2017). The context is not so relevant here but the work is what I call a sound frame. It’s a metal frame the size of a large painting hanging 30cm from the wall. The rear side contains speakers, reproducing a two-dimensional auditory image (in loop) of the process of drawing on paper (Figure 2). I wasn’t really aware of spatial sound capturing at that moment and empirically designed this system, which performed well.

This artwork was the beginning of my journey into spatial sound exploration, resulting in installations (Radison – Kindergarten, Neuron – Choral Flow, …) to Ambisonics concerts (Ethereal Threads, Nanotonal Ascension for 24 Voices, …) and VR compositions (Polytox, Boredom, …), concepts like that of the iPixel and the idea to produce AcouStasis, a wavefield wallpaper containing arrays with hundreds of electrostatic speakers to create instant acoustical emulations, eliminate noise or play the spatial sound of a rainforest in any room (Figure 3).

Figure 4 — Possible shapes of 3D sound objects.

A walk through Sound Fields

With a background in 3D modelling and animation I soon realised the unexplored possibilities of 3D sound experiments and compositions I heard in the Wavefield lab at IPEM. I saw opportunities in adopting and extrapolating 3D modelling and animation techniques from visual practice to the auditory domain. I quickly got stuck on the basics: what is a 3D sound object —in VR terms—? How big can it be, what shape can it have (Figure 4), does it has a surface, what’s inside it…? And even further: besides being volumetric, could it be fog-like, a particle system or NURBS-surface and what about reflection and absorption of multiple sound objects (including walls and people) in a space, (trans)-morphing sound objects and more…


Taming the octopus was key and I started by studying 3D recording techniques. Almost all of them are space-oriented (covered in the next chapter) and one of them caught my attention: the Soundfield microphone invented in 1978 by Michael Gerzon and Peter Craven. Their theoretical design was applied into a microphone system by Calrec Audio Limited, who launched the Soundfield microphone. It has four closely spaced cardioid microphone capsules arranged in a tetrahedron, facing outward. A full-sphere surround sound format was later defined by a group of researchers in the UK and named Ambisonics4.

The number and types of microphones are defined for different “resolutions”, known as first order Ambisonics (FOA) when four microphones are used or higher order Ambisonics (HOA: 2nd, 3rd, …) when additional microphones are added. This brought me back to the sonogrammetry device. Why not use the tetrahedron model of a Soundfield microphone and invert it, changing the POV (point of listening) on the "other side" of the wavefield. Space-oriented becomes object-oriented (Figure 8). The maths was already there. The notation of this “inverted” FOA wavefield capture becomes ~FOA.

Figure 5 — An isocahedron setup in an in an anechoic chamber.
Figure 6 — Surrounding 64-channel spherical microphone array.
Figure 7 — Plenacoustic analysis.

A Brief Overview of Research on 3D Sound Recording

Common more complex (i.e. not mono or stereo) techniques for 3D recording using multiple microphones directed to the sound source are mostly orchestral or multi-instrument (ensemble) setups. They may look similar to IrSOb but are space oriented recordings.

The 22.2 multichannel microphone setup (Will Howie et al. 2016) is an example of such a very complex setup with discrete channels and different flavours of omni- and cardioid microphones. A wide range of complex setups can be found for surround recordings among them the popular ORTF-3D using 4 cardioid and 4 super-cardioid mics, the Hamasaki Cube with 4 figure-8 mics, the OCT-3D 2L-Cube and the Decca Cuboid to name some.

The 42-channel spherical microphone array is a human sized icosahedron for measuring acoustic violin properties (National Institute of Technology, Japan — Figure 5)

And then there is Markus Noisternig, Franz Zotter, et al. with their surrounding 64-channel spherical microphone array. It's set up at IEM for acoustic research in a semi-anechoic chamber (Figure 6).

Antonio Canclini et al. at the Politecnico di Milano describe a method for the estimation of the three-dimensional radiation pattern of violins, during the performance of a musician. A plenacoustic microphone array captures the energy radiated by the violin in different directions using beamforming based on sub-arrays (Figure 7).

All methods mentioned above (there are many more) have their own specific advantages and uses. The more complex ones are only possible in lab or equipped stage environments. Though the ideal IrSOb recording would also be in an anechoic room for accurate archiving purposes, it offers proficient results for the reproduction in a 3D VR experience in any environment (also outside).

Figure 8 — Left: capturing spatial wavefield sound. Right: capturing object wavefield  sound.

IrSOb The Iridescent Sound Object

When recording with a SoundField microphone, the direct output from the 4 cardioid microphone capsules is called the A-format (figure 8, left)5. The Ambisonics B-format is the standard audio format produced from the A-format after the Soundfield calculation, consisting of W (pressure signal), X (front-back information), Y (left-right) and Z (up-down information). The capsules are closely positioned in a tetrahedron pointing outward and individually identified as FLU (front left up), FRD (front right down), BLD (back left down) and BRU (back right up). In itself, the signals are unusable since they only contain phase-related information instead of channel based information.

For the IrSOb model, this design is inverted. The four capsules of the capturing apparatus aren’t at the centre anymore but have been moved to a sphere around the point of interest (object K), now at the centre. The cardioid microphones pointing inward are still arranged as a precise tetrahedron to avoid phase shift and to eliminate time of arrival calculation. The virtual listener (participant P) can be anywhere around (or in) the virtual 3D object. The ~A-format (Figure 8, right) allows for the calculation/reconstruction of the 3D sound of the object in the VR space.

The IrSOb model has three zones, each having a set of 4 sound channels maximum:

  • the O-zone (object zone): 4 channels to contain the direct object sound (the ~A-format);
  • the A-zone (ambient zone): 1, 2 or 4 channels for the reflected space sound (reverb);
  • the I-zone (inside-zone): 1, 2 or 4 channels for the object’s internal sound.
Figure 9 — Representation of the inverted tetrahedron setup and the three zones of the IrSOb concept. K: sound object, P: participant, E: Event Horizon.
Figure 10 — First test recording setup. L: the sound object is temporarily replaced by
 an Ambisonics mic recording the "room tone". TR: the sound object with distinct front/rear sources. BR: 
accurate distance and angle measuring device.
Figure 11 — The IrSOb M4L device is connected to the HybridEngine at IPEM.

The Prototype

I developed two IrSOb proof-of-concept systems. The first at IPEM’s ASIL6 3D VR sound lab as an E4L7 and Max for Live device that I wrote and the second was Earis, the extra-lab binaural headset device.

Both systems translate the ~A-format into a two-dimensional horizontal plane (the I-format),resulting in a 360º radial sound image of the object experienced in a headphone.

The 3D sound image of the object is called the ~B format. A two-channel orthogonal raw left/right signal is derived from the I-format and further processed in function of the location and orientation of participant P regarding to the virtual source K. This processing results in the O-zone and includes:

  • Sound level and frequency filtering (psychoacoustic adjustment8) in function of the angular position for the left and right ear of the participant regarding to the VR sound object (a basic binaural filter);
  • the stereo separation (field width) and volume level in function of the distance to the VR object (closer to object results in wider stereo field);
  • the balance of direct (O-Zone) and ambient (or reflected, reverb – A-zone) signal in function of distance and head rotation.

The source of the reverb signal can either be simulated (e.g. the Convolution Reverb in Envelop for Live) or recorded by an extra set of microphones (same axis of the direct mics, pointing 180º away from the object) to create the A-zone. Obviously, the use of a Soundfield mic is difficult if not impossible since it should be placed in the exact centre of the tetrahedron where the object is9. The use of four extra mics results in a space oriented reverb image, that will equally change in function of position and listening angle. It can be treated very similar as a B-format.

The Event Horizon E, is the spherical border between the O-Zone and the A-Zone (Figure 9). When crossed, the participant “enters” the object. Since the IrSOb sound object has a (virtual) volume, a third set of mics can capture sound inside the object (the I-zone) and will vary according to the object or desired effect. A single mic in a cello could suffice, contact mics on a skull, a Soundfield mic in a car, a hydrophone in an aquarium… The use of MEMs mics can extend this interior sound experience considerably. The content (as with the whole IrSOb object) of this peculiar zone can be anything from a real recording to a fantasised inner world sound montage (in 3D).

The IrSOb model is technology independent. The first test started with recording a sound source with distinct spatial sound radiation (Figure 10). Then followed the development at the ASIL lab. I patched a MAX for Live device (Figure 11) that connects to the lab’s HybridEngine M4L device in order to get the MoCap data (position in the space and head rotation) of a participant and calculate the binaural signal related to the virtual object positioned in the lab VR space. The processed signal was then transmitted from the central server over analog UHF to the headphone of the participant. The (quick and incomplete) overview (Table 1) below gives an idea of the difference between a lab environment and the developed Earis hardware experience (both using the same IrSOb model).

One advantage of the IrSOb concept is the independent recording (or synthesising) and re production technique. One can record sound in 3D in a different order than that of the reproduction system, making it flexible and future proof.

Property

Lab

Earis

Number of simultaneous users

At the time of testing: two. Restricted to central processing, MoCap simultaneous users and wireless transmission channels. (Qualisys max. 10)

As stated by the manufacturer: up to 8 beacons and 64 users.

Number of simultaneous IrSOb objects

(not tested)

Depends on number of users (of the MoCap and the central processing of the data per user to be streamed per user.

Depends on the participant’s processing unit (CHAP) and data communication over the UWB and WiFi.

Latency

Depends on number of users of the MoCap and the central processing of the data per user to be streamed per user.

Less influenced by number of users and beacons.
Dependent of CHAP processing unit.

Accuracy

Centimeters

Centimeters (more beacons increase resolution but add to latency)

Location

Constraint to lab

Everywhere

Table 1 — Comparison of lab and field experience.

Figure 12 — Multi-user, multi-object.

Multi-user, multi-object

The Earis system supports multiple simultaneous users. The maximum users on one system with several beacons depends on the processing power of the user devices and the radio traffic of the UWB beacons and WiFi (Table 1). Several IrSOb virtual objects can be positioned in the VR sound space (Figure 12). The position of the users in the space is known and this information can be used to dampen the sound coming from the object in regard to the position of a specific user. In the example below, an object A propagating its sound towards a user P1 would be perceived directly (1). This user would block part of the sound travelling (2) to user P2 who also hears the sound of object A (3) reflected by object B (4) and that of B directly (5). User P3 hears the partially blocked sound of A (6) and not the faint sound of B (7), the latter ”to weak” to reach P3. Reflection and absorption from walls and fixed elements (8) as well as users could also be accounted for but are not included in the current version. Adding to the calculation load are: the number of beacons, simultaneous users, IrSOb objects, reflections/absorption (by object, user and VR walls).


Figure 13 — An open headphone with the prototype hardware mounted on top.
Figure 15 — Diagram showing the architecture of the Earis hardware.

The Earis Headset

The Earis system consists of a semi-open headphone set to reproduce the processed IrSOb audio channels in a binaural signal. The head position and orientation tracking unit (HPOT) is mounted on the headphones (Figure 13). The input from the location and orientation sensors (Figure 14) is processed in useful data sent to the channel processor unit (CHAP). Here the audio processing is done: the calculation and mix of the three audio channel sets O, A and I resulting in a binaural audio signal, sent to the headphone (Figure 15).

At the time of construction, the CHAP is a RaspberryPi 4, running Python code (Figure 16) to process eight audio channels stored on the unit’s µSD card. The limitation on channels for the PoC (8 instead of the 12 specified in the IrSOb model) is a restriction by design, imposed by processing power limits. A 3300mAh LiPo battery provides enough energy for 5 hours of operation. The UWB beacons (Figure 17) are also ESP32 boards with LiPo batteries.

The HPOT is a C++ coded ESP32 based microcontroller board with WiFi and UWB radio transceivers for the location tracking and calculation. Connected to it is an inertial measurement unit (IMU) for measuring the orientation resulting in an Euler Vector and the iRose, a compass-like hub containing 8 IR receiver LEDs (Figure 18). They pick up a digitally coded signal sent from an omnidirectional IR emitter (Figure 19) for real-time drift correction of the IMU (due to magnetic interference inside a building).

The core mechanism for accurate centimetre-level location tracking of a participant P (depending on the number of anchors) is the HPOT’s onboard DecaWave (Qorvo) ultra-wide band (UWB) wireless radio transceiver operating at 5–9GHz and running a communication protocol used for trilateration10. For this, (at least) two fixed UWB beacons (A1, A2) are required with known position and the UWB tag of the participant as the third unit (Figure 20). The IrSOb VR object K has a known virtual position (would the object also move in the space during a session, additional calculations would be necessary). This UWB technique provides precise indoor location tracking with an accuracy of 5cm maximum, depending on the number of units (more units mean higher and more redundant data), (wall) reflections, (participant) damping and radio interference.

Figs 14 16 17 18 19
Figure 20 — The trilateration process. A1, A2: beacons. K: IrSOb VR object, P: Participant, r radius P—K,  phi: relative angle P to K in the space, theta: relative listening angle P.
Figure 21  — The iRose hub containing 8 IR receiver/decoders with precise angle light slots.
Figure 22 — The drift correction paradigm: a segmented phase lock system of resolution of n/2/2.  IR beacon I is positioned at object K. The IMU reading θA of the head H is off-track according to the IR beacon detected by p1.

The iRose

While the UWB setup takes care of the accurate position of the participant, we also need head tracking for a complete spatial orientation. The latter is a more troublesome affair. Head direction is monitored by a 9DoF11 inertial measurement unit (IMU) in 6-axis mode. The internal magnetometer is turned off to avoid heading drift or sudden jumps because its onboard fusion algorithm auto-calibrates the internal magnetometer, which is easily affected by nearby magnetic and electrical interference. Nevertheless, drift still occurs. Without a magnetic reference, the sensor relies solely on the accelerometer and gyroscope to track movement. This setup causes the heading to drift continuously over time, while roll and pitch remain stable.

To correct this drift I developed a real-time angular realignment device. The iRose provides continuous (re-)calibration of the drifting IMU sensor data. It is built around a central omnidirectional infrared emitter (beacon). The IR beacon (visible in Figure 1) sends a specific code in bursts at 38kHz. It can be mounted near the ceiling above the virtual sound object to avoid blocking of the signal. Other positions are possible and the use of several beacons is also an option to increase resolution and create redundancy. Caution is needed to avoid IR reflections (windows, sleek surfaces) but an extended algorithm could get rid of most of that. The built iRose receiver is a compass-like hub containing eight IR receiver LED’s (as in a wind rose). The case holding the compass PCB12 and the IR LEDs is designed with eight “light gates” with a 45° entry slot for each IR receiver (Figure 21). When combining the reception of 2 adjacent LEDs a doubling of the angular resolution is possible: 22.5°. A 32 channel iRose would have an angular resolution of 5.625° which is sufficient for a most applications.

When an IR code is received by a LED in the hub, the code holds the angle of correction. If the drift of the IMU gyro exceeds 2π/8/2 (the basic resolution of 45° with 8 IR receivers, in both directions (/2)— then it is reset to the correct angle. E.g.: LED E (east or 90°) detects the calibration code and the IMU Gyro angle says 95°, then nothing happens (|95 - 90| > 22.5). Should it be for instance 65° then a reset to 90° will happen (|65 - 90| > 22.5 — Figure 22). This, together with a slew factor avoids constant erratic jumping of the orientation13.

The slew factor mentioned above goes mostly unnoticed since it matches the system’s latency due to processors and software bottleneck handling the sensor data and even more, the audio channel processing. Faster processors (RPI5…) and preemptive maths could bring this to the unnoticeable. On the notion of embodiment it is worth mentioning that the latency (the lag of the audio signal when a participant is moving) results in the person moving slower and observing more focused.

IrSOb, the (l)on(e)ly 3D VR sound object?

Current virtual reality development tools such as Google VR’s GVRAudioEngine and Unity’s Audio Spatializers do not provide 3D VR sound objects14. What they offer is spatialisation by modifying mono omnidirectional audio clips so they appear to come from a specific direction, angle, and distance relative to the listener. They achieve this by parameterising attenuation (decreasing the sound volume as the user moves further away from the virtual object) and room acoustics (simulating early reflections, late reverberations, and occlusion based on virtual walls and geometry). Except for the ear-to-ear time delays based on the relative positions and orientations, IrSOb simulates real-world acoustics by modifying sound volume, frequency filtering and correction for binaural listening, starting from a 3D VR sound object.

Creative applications

The Iridescent Sound Object model extends the closeness/distance dynamic in film and sound productions. As a visual storytelling tool, camera proximity and the use of different focal points15 control the emotional bond between the audience, characters, and their environments. The same applies for this dynamic in audio that relies on psychoacoustics, frequency balance, and spatial cues. When a sound source moves closer, it gains volume, high-frequency detail, transient punch, stereo spreading and a low-end boost known as the proximity effect.

The possibilities with IrSOb are stretching from obvious dramatic examples: —inside/outside a submarine cruising through a coral reef, or more scientific experiments like capturing a cello played live where the researcher or instrument builder can observe the sound from all directions and even inside the cello or the musician’s head.

Or think of it as a new strong narrative element, where in a dialogue or even silent scene, the outside is the sound of what we hear (spoken word, …) and the inside is that person’s thoughts, a key missing ingredient when adapting books to film.

Many more examples of recording situations can be given. One application discussed with Valeria Mignaco, docArtes researcher at Leiden University was to capture the intimate household sphere where female amateur singers in Early Modern England sang a repertoire of devotional songs. This is a kind of performance that can not be experienced as intended on a stage. As she quotes:

“I was intrigued by Philippe’s presentation at the Electronic Luthérie conference and experienced the IrSOb demonstration. With the headphones on, I intuitively moved through the space, following the sound. Soon, I became aware of a sensation of moving around the complex VR object, experiencing it from outside, beside and ultimately even within the sonic object, while remaining aware of its physical boundaries. This was precisely what I was looking for in my research: the experience of hearing a musical performance from the physical location of the artist.”

Stepping aside from recorded or even live applications, a purely synthesised object can be created, radiating different sounds, timbres or effects from all directions. As the participant walks around or in the object, the sound changes in function of their position in relation to the object. Of course, the Iridescent Sound Object can move in the VR space as well. Interactive parameters could be added to control the size, position and content of the object adding to even more possibilities of unheard experience.

From VR to HO

The three basic layers (A/O/I) of IrSOb —intended as a tool for realistic sound recreation) and its angular control over spatial content can be expanded to several layers, creating an onion structure (or spheric 3D matrix) controlling the sound. When taking this even further and increasing the resolution substantially, it would become a holographic object (HO) or Wavefield sphere in itself.

Image 23 — The recording in IrSOb of the press at the Industriemuseum (only 2 mics are visible).

Case Study / PoC: The Heidelberg Press Sound Project

The Heidelberg Press Sound Project started as a case study for the IrSOb model but goes beyond that. The 3D sound recording is that of the complex and sophisticated mechanism of a historic paper press while it is printing the message “This is Your Page Between the Street and the Beach Beneath”. One hundred copies were printed with a historical Schnellpressenfabrik Heidelberg platen press at the Industriemuseum in Ghent (Figure 23). The numbered and signed prints are exhibited as a one meter high stack pile, allowing people to walk around and listen to it while being printed, in IrSOb VR sound.

The four cardioid microphones are meticulously positioned in a tetrahedron. A precise distance and angle is necessary to recreate the correct image. Extra shielding is attached on the back of the mics to avoid inevitable reverb pickup.

The sentence on the page is an extension of the original “The Beach Beneath the Street” expression that originated within the Situationist International movement, an organisation of social revolutionaries made up of avant-garde artists, intellectuals, and political theorists. Among them: Guy Debord (FR –The Society of the Spectacle) and Raoul Vaneigem (BE – The Revolution of Everyday Life). At that time, the distribution of printed pamphlets was the best way to propagate visions and announce events.

During the Electronic Luthérie conference that took place at Orpheus Institute in January 2026, attendees could try the IrSOb v3.5 edition with the The Heidelberg Press Sound (Figure 1).

Conclusion

Though the IrSOb model, its use in a Wavefield lab and the Earis technology have proven to deliver a new kind of method and experience, this is only the start. As an open source project (Github link below) it allows to be tested, dissected, improved, altered, enhanced and further developed into specific systems or as a standard model. It could start with:

  • a hardware upgrade to a single device with more processing power;
  • a 16 channel iRose drift correction and statistical algorithm to remove sensor noise;
  • the possibility to add more IR beacons and choose location;
  • auto-calibration of everything at setup;
  • addition of more tags;
  • a remote UI to set all parameters and a CMS to add objects, sounds, behaviour…
  • real HRTF implementation;
  • live streaming input;
  • define the metadata and methodology and a 3D notation for space-related scores;
  • a 2nd order version (3 x 8 channels)…

Some may question if IrSOb is an instrument at all. Wikipedia learns us that “A musical instrument is a device created or adapted to make musical sounds. In principle, any object that produces sound can be considered a musical instrument— it is through purpose that the object becomes a musical instrument. A person who plays a musical instrument is known as an instrumentalist.” IrSOb may seem an audio recording method at first sight, but the manipulation of the hardware or in a VR environment enables the beholder to become an instrumentalist, manipulating the sound outcome. Even more, when creating a VR 3D sound object from scratch, we build a VR instrument with the purpose of producing sound through interaction. I hereby invite artists, engineers, VR sound artists, composers and researchers in a call to explore and expand IrSOb.

To conclude my PoC I am creating a simple but colourful synthetic IrSOb sound bubble comparable to that of a an iridescent soap bubble —from which the model takes its name—, translating the visual effect to that of the sonic realm. Currently I’m also working on a synthesised (lab created) IrSOb object for an interactive composition called Tinnitude. This is a related study on tinnitus within my PhD, a filed in which I have previously been doing research on and created specific tinnitus-related works.

Acknowledgment

Thanks to my promotors at IPEM, Dr. Prof. Pieter-Jan Maes for the moral support and use of the ASIL lab and Dr. Bart Moens for sharing his expertise and the willing technical support of the Qualisys MoCap and HybrisEngine systems.

Thanks to my friend Bechara Yared, CEO of Look, a company making high quality and interesting museum experience devices who believes in the potential of IrSOb for museums and heritage applications.

Thanks for the kind cooperation of Marie Kympers (IM), Jean-Pierre Berth (press operation) and Armina Ghazaryan (type setting) for the printing and recording process of the Heidelberg press at the Industriemuseum.

Photo and illustration credits

All photos and illustrations by the author, except:
Figure 5 Yuya Nishimura.
Figure 6 Markus Noisternig.
Figure 7 Lily M. Wang.

References

Bachelard, Gaston. The Poetics of Space. Translated by Maria Jolas. Boston: Beacon Press, 1964. Originally published as La Poétique de l’espace, 1958.

Blauert, Jens. The Technology of Binaural Listening. Berlin: Springer, 2013.

Canclini, Andrea, Luca Mucci, et al. “A Methodology for Estimating the Radiation Pattern of a Violin in a Plenacoustic Setup.” In Proceedings of the European Signal Processing Conference (EUSIPCO), 2561–65. IEEE, 2015.

Gerzon, Michael A. “The Design of Precisely Coincident Microphone Arrays for Stereo and Surround Sound.” Preprint L-20, 50th Convention of the Audio Engineering Society, London, 1975.

Hemmens, Alastair, and Gabriel Zacarias, eds. The Situationist International: A Critical Handbook. 1st ed. London: Pluto Press, 2020. https://doi.org/10.2307/j.ctvzsmdw0.

Howie, W., T. Shono, Y. Iwaya, and C. Kyriakakis. “A 22.2 Multichannel Microphone Array for Immersive Audio.” Paper presented at the 141st Convention of the Audio Engineering Society, 2016.

Katz, Brian F. G., and Piotr Majdak, eds. Advances in Fundamental and Applied Research on Spatial Audio. London: IntechOpen, 2022.

Martin, G., N. Holighaus, and P. Balazs. “Subjective and Objective Evaluation of 9-Channel Three-Dimensional Acoustic Music Recording Techniques.” Paper presented at the 141st Convention of the Audio Engineering Society, 2016.

Nishimura, Y., N. Yasui, et al. “A Consideration on the Sound Radiation Pattern of Violin.” International Journal of Emerging Engineering Research and Technology, 2016.

Noisternig, Martin, Franz Zotter, and Alois Sontacchi. “3D Binaural Sound Reproduction Using a 64-Channel Spherical Microphone Array.” Institute of Electronic Music and Acoustics, University of Music and Performing Arts Graz, 2011.

Pfanzagl-Cardone, Elisabeth. The Art and Science of 3D Recording. Cham: Springer, 2023.

Wang, L. M., and C. B. Burroughs. “Directivity Patterns of Acoustic Radiation from Bowed Violins.” The Journal of the Acoustical Society of America 106, no. 6 (1999): 3165–76.

Zhang, R., R. Meng, et al. “Modelling Individual Head-Related Transfer Function (HRTF) Based on Anthropometric Parameters and Generic HRTF Amplitudes.” CAAI Transactions on Intelligence Technology 8, no. 2 (2023): 364–78.

Zotter, Franz, and Matthias Frank. Ambisonics: A Practical 3D Audio Theory for Recording, Studio Production, Sound Reinforcement, and Virtual Reality. Cham: Springer, 2019.

Research Projects and Online Resources

“John Dowland’s A Pilgrimes Solace: Devotional Songs for Female Amateur Singers in Early Modern England.” docARTES. Accessed August 14, 2026. https://www.docartes.be/en/research-projects/john-dowlands-a-pilgrimes-solace-devotional-songs-for-female-amateur-singers-in-early-modern-england.

Ritsch, Winfried. “IEM Atelier Algorythmics—ACRE-AMB, Klangdom.” Algorythmics. https://algo.mur.at/ritsch.

Druez, Philippe. “Related Artworks and Research.” http://philippedruez.be.

Druez, Philippe. “IrSOb: The Iridescent Sound Object—Code and Diagrams.” GitHub. https://github.com/HiPhiPhD/IrSOb-The-Iridescent-Sound-Object.

Imprint

Issue
#8
Date
07 September 2026
Category
Review status
Anonymous peer review

Footnotes

  • 1 1. Dolby Atmos surround sound, XPERI’s DTS Headphone:X, Creative Labs’ Super X-Fi, Windows 10 Sonic, Apple Spatial Audio, Sony 360 Reality Audio (WalkMix) and Waves Nx (combining room emulation and motion tracking)
  • 2 2. Other providers of motion capture devices and software like OptiTrack, Rokoko and Move are equally usable. 
  • 3 3. DAWs (also called SAWs “spatial audio workstations") like Steinberg Nuendo / Cubase and DaVinci Resolve / Fairlight and others…
  • 4 4. There is a comprehensive work by Franz Zotter and Matthias Frank on Ambisonics (see references) for those who want to dig deeper into the matter.
  • 5 5. Some figures may show inverted as a macron diacritic above the letters instead of prepended by a tilde but they both denote “inverted”.
  • 6 6. ASIL: the Art & Science Interaction Lab offers an acoustically treated black-box with
    Qualisys Motion Capture and an 80 speaker audio system for spatial & immersive audio.
  • 7 7. E4L (Envelop for Live) is an open-source spatial audio production toolkit allowing artists to perform, produce and explore DIY 3D VR sound projects. E4L operates within Ableton Live Suite.
  • 8 8. No real psychoacoustic or phase related adjustments are done for now. A true HRTF filter could be applied in a later stage. A deeper understanding of binaural listening can be read in the work of Jens Blauert's The Technology of Binaural Listening.
  • 9 9. An off-centre recording is acceptable in some cases and/or a mathematical correction is possible.
  • 10 10. Trilateration is the geometry that uses distances to calculate angles between minimum three objects with known position (triangulation measures angles). The distance is calculated from the ToF (time of flight) between units send and receive. It is used by GPS units tracking the position of multiple satellites (multilateration).
  • 11 11. 9DoF — Nine Degrees of Freedom, refers to an IMU that tracks movement across nine individual axes using three distinct 3-axis sensors: an accelerometer, a gyroscope, and a magnetometer. It measures linear acceleration, rotational velocity, and magnetic heading to determine spatial orientation. The used IMU is a the Bosch BNO055.
  • 12 12. Due to a lack of pins on the ESP32 dev board I could only use 4 of the 8 LEDs resulting in a 90º (or interpolated 45º) resolution.
  • 13 13. In this dynamic system a Kalman Filtering (Linear Quadratic Estimation) or Bayesian Inference Filtering could reduce erratic sensor measurements considerably without requiring extra hardware.
  • 14 14. hough you can use a stereo sound source with Unity’s VR Audio Spatializer, it is generally not recommended.
  • 15 15. A close-up with a 35mm lens may frame the same as with 100mm when the camera is further away but the latter loses the perspective information hinting of closeness.

Leave reply

Your browser does not meet the minimum requirements to view this website. The browsers listed below are compatible. If you do not have any of these browsers, click on the icon to download the desired browser.