Abstract
Diocletian’s palace with its cellars represents one of the most important cultural heritage sites of the ancient Roman civilisation on the present-day Croatian territory. The cellar complex has been rediscovered only recently and has been preserved remarkably well due to its centuries-long concealment beneath mediaeval urban matrices. An archaeoacoustic analysis was performed on a selected single-nave hall as a small part of this complex. A model of the hall was developed in room acoustics simulation software and calibrated based on the results of field measurements. Acoustic suitability of the hall for speech-based events and music performances was then evaluated according to contemporary objective criteria, and the findings were compared with the results of similar studies performed on other heritage sites. The hall was found to be very well suited for speech in terms of intelligibility and mid-frequency reverberation, thus showing potential for revitalisation, with excessive low-frequency reverberation in the hall and reduced audibility in the farthest part of the audience as potential issues. With a feasible audience size, the hall is not reverberant enough for music performances but provides high clarity. In terms of sound strength, the hall is suitable for solo performers or small ensembles. Excessive perceptive broadening of the sound source is expected due to strong early lateral energy. In terms of traditional Dalmatian a cappella singing, the acoustics of the hall are likely to support and enhance such performances.
1. Introduction
In the context of ancient Greek and Roman civilisations as the foundations of modern Western civilisation, the writings of Vitruvius and the recorded development of ancient theatre design suggest that Greeks and Romans possessed an understanding of architectural geometry, proportions, and materials and the influence they have on the projection of sound and the overall acoustic perception of a space [1,2,3,4,5]. Practice-driven preferences were shown through different typology-based scales: odea for music performance, theatres for drama performance, and amphitheatres for gladiatorial combat, displaying the evolution of acoustic requirements for spaces based on their function [3,5,6,7,8]. The available literature suggests that ancient Greek and Roman theatres were designed with a strong emphasis on speech intelligibility, audience coverage, and tonal character, as they were built for drama performances and public speech [3,9,10,11,12,13]. To this day, these ancient venues are still considered excellent performance venues [3]. Yet, historically used intuitive and perceptual guidelines for acoustic quality were likely different from the objective guidelines and criteria used nowadays, as substantiated by various studies of Roman heritage sites [3,12,13,14,15], which yielded the values of objective parameters outside the typical ranges defined by the ISO 3382-1 standard [16].
The research on the acoustics of heritage sites has gained momentum in recent years. However, the interest is spread quite unevenly between outdoor and indoor ancient sites [1,3,9,10,11,12,13,14,17,18], as well as between sacred and profane spaces from later periods [6,7,19,20,21,22,23,24]. Although most of the examined ancient heritage sites were outdoor venues, predominantly open-air Greek and Roman theatres and amphitheatres, the Ten Books of Architecture by Vitruvius [5] and later research based on inherited Roman spaces defined certain architectural and acoustics-related preferences for indoor spaces as well. In Book 6.3 [5], Vitruvius sets out proportional rules for both square and oblong spaces used as dining rooms and in domestic interiors in general. These rules suggest the awareness of the influence of room geometry and auditory function on room acoustics [5], while indicating that Roman architecture treated sound as an active dimension of spatial experience rather than as a secondary by-product of form [3,4,9,25,26]. Regarding the guidelines related to acoustic treatment of spaces, in Book 5.2, Vitruvius states “the inside walls should be girdled, at a point halfway up their height, with coronae made of woodwork or of stucco” to maintain speech intelligibility [5]. Later in Book 5.3, he also explains the importance of the usage of various materials, such as wood, stone, ceramics and bronze, in gaining the preferred tonality, thus enriching the oratory and rhetoric “with greater clearness and sweetness to the ears of the audience” [5]. Lastly, the use of solid materials such as masonry, stone, or marble is mentioned as predominantly reflecting low-frequency energy back into the venue, thereby enriching the acoustic environment with a warm and enveloping character [3,4,5]. In light of this, contemporary criteria such as ISO 3382-1 metrics [16] certainly remain a valuable descriptive tool in heritage research, as a clear relationship has been established between objective parameters and the corresponding perceptual dimensions such as reverberance, clarity, loudness, envelopment, perceived source width, speech intelligibility, and others [27,28]. However, to establish and fully understand the relationship between the architecture, acoustics, and function of ancient Greek and Roman heritage sites, such a tool used as the sole basis for evaluation might not be sufficient [2,7,10,17,20,21]. This opens the door for perception-based evaluation as a valuable complement to purely objective evaluation and a potential tool for gaining a comprehensive understanding of cultural perception preferences in terms of the acoustics of ancient heritage sites. Overall, indoor and residential spaces from the Roman era remain underexplored, especially the ones preserved in a near-intact condition. In addition, many of these sites are assessed in terms of archaeological value only, while aspects of intangible heritage remain vague at best.
As a part of the efforts focused on developing the field of archaeoacoustics within the Croatian research community, it is worthwhile to understand the specific circumstances and history of the present-day territory of Croatia. Throughout recorded history, this territory belonged to and was inhabited by different peoples, cultures, and societies, all of whom left their mark. A fortunate circumstance is that the cultural heritage sites on the territory of Croatia are generally well preserved. In that light, Croatia constitutes an especially promising ground for future archaeoacoustic studies, while research among the Croatian academic community in this field is still in its infancy.
This study represents the initial stage of broader research, focusing on investigating a single-nave hall as a part of the cellar complex of Diocletian’s palace located in the present-day city of Split, Croatia. The entire site represents one of the most important historic urban sites on the Dalmatian coast. Having functioned both as a palace and a castrum, it is a unique example of a building erected to establish and support Roman territorial control over the Dalmatian region. Over time, the palace and its surroundings have ultimately developed into the city of Split [29,30]. The palace was built as the residence of emperor Diocletian, to be used after his abdication in 305 AD. The location of the palace at the centre point of Dalmatia was presumably chosen to represent the power of the emperor. Given the status of its occupant, it is to be presumed that the palace was built with the highest regard for architecture, design and function.
While the palace complex can be regarded as an example of domestic architecture at the level of sensory control and socially structured use of space, a direct comparison to better-known typologies dominating in archaeoacoustic literature proves to be quite difficult. To address this gap, the studied hall has been chosen for an archaeoacoustic analysis as a unique space that differs from other heritage sites of the Roman and subsequent eras. In contrast with open-air theatres and amphitheatres as outdoor venues, this hall is a fully enclosed indoor venue and part of a coherent substructure of a larger and highly elaborated complex. The difference to other indoor spaces of the era, such as catacombs, cryptoporticus, and hypogea, is reflected in size and shape, state of preservation, connection with adjacent spaces, and material used in construction [31,32]. The distinction between the hall and spaces such as basilicas, baths and large civic halls lies in its more compact volume and outdoor coupling due to the presence of clerestories and open façades [6,18,21,33]. Finally, the difference between the hall and the later basilica-derived forms, namely early Christian single-nave churches, is reflected in their function. While the exact function and purpose of the investigated hall remain unknown, the main focus in this type of building was put on music and liturgical chant [6,7,19,20,22,23].
The second reason to choose this particular hall as an object of investigation is its almost pristine condition, as it is a rare example of an indoor space from the Roman era to possess this quality. The entire cellar complex was essentially buried, forgotten, and embedded into mediaeval urban fabric and infrastructure and was excavated only seventy years ago [34]. As a consequence, the original state was quite well preserved, and the damage that occurred over time is attributed to the historic practice of using the cellar complex as a destructive sewage system, dating back to the Middle Ages [30]. In addition, other than the necessary repairs of the floor, no architectural intervention of any kind was designed or implemented in the hall. As the entire complex is a World Heritage Site protected by UNESCO [35], it is unlikely that any intrusive interventions will be permitted in the future.
The research conducted in this study is focused on (1) recreating the acoustic situation in the hall via field measurements and subsequent modelling and calibration and (2) performing an evaluation of the hall based on contemporary acoustic criteria in terms of its acoustic suitability as a potential venue for speech-based events and music performances. In the latter stage, particular attention is given to reverberation time, clarity, speech intelligibility, sound strength, and early lateral energy fraction as the contemporary objective criteria chosen for acoustic evaluation.
At present, (1) the true meaning of the perceptual criteria used by the Romans is uncertain in terms of modern understanding of acoustics, (2) there is no viable way to assess the damage to the hall caused by centuries of use as a sewage system and no solid data on its original design, and (3) there is no solid data that would indicate the original function of the hall. In light of this, the study does not attempt to (1) assess the acoustics of the hall in terms of Roman acoustic preferences, (2) recreate the original state of the hall, or (3) speculate on the exact original function of the hall.
By conducting this study, the authors hope to contribute to the field by (1) documenting and performing an archaeoacoustic analysis of a space that is unique in terms of architectural and historical context, as well as the state of preservation, and quite different from the ones predominantly studied in terms of archaeoacoustic analysis, and (2) providing a foundation for the revitalisation of the hall in a contemporary context by assessing its potential for adaptive reuse for both speech-based events and music performances in terms of acoustic suitability.
2. Methodology
2.1. Architectural and Historical Overview
Historically, Diocletian’s palace is a unique example of cultural heritage in terms of architecture and urban planning. The duality of its function is also visible in the inner layout, as it consists of the two main levels: the ground floor organised around the Cardo and the Decumanus, and the cellars located beneath the southern part of the Peristyle (a forum on a castrum scale) [36], as shown in Figure 1. While the ground floor was reserved for imperial quarters, the cellars were devised as a functional and structural counterpart, rather than a space of representation [30,34,37].
Figure 1.
Floor plans of Diocletian’s palace; preserved parts [34].
The palace was constructed at the turn of the third to the fourth century AD by the Roman emperor Gaius Valerius Aurelius Diocletian to serve as his residence after his abdication in 305 AD, but it was abandoned soon after due to the invasions of barbarian peoples, the rise of early Christians, and finally the fall of the Roman Empire [29]. After the fall and during the Middle Ages, the outline of the ground floor was completely altered. The original of the palace layout can be partially reconstructed by mirroring the layout of the auxiliary spaces beneath, confirming the primary role of the cellar complex as a supporting structural system. The cellar complex had additional infrastructural and sanitary functions, as the cellars served as cesspits or sewage-collection chambers for the mediaeval city built on top, i.e., at ground floor level [29,38]. Being forgotten and serving as a part of a sewage system paradoxically protected the geometry and material integrity of the entire site until the excavation of the substructures, conducted between the 1950s and 1970s. The original state of the cellars was documented after removing the infill, revealing construction techniques and engineering principles typical for the late period of the Roman Empire [29,30,38]. Their main features are characterised by massive stone masonry, regular ashlar courses, opus mixtum (lower quality crushed stone with four rows of bricks approximately every 150 cm in height) construction techniques, barrel vaults and arches, etc. [30,34]. The spaces of the cellar complex were built according to similar construction principles, forming a sequence of mid-sized rectangular halls that articulate the coherent basement complex [29,34,37,38]. Due to the high amount of humidity and low ventilation of the space, a thick layer of porous, nowadays partially damaged tufa has formed on the ceilings over time as another unique aspect of the site.
Located in the southwestern part of the cellars (see Figure 1), the investigated single-nave hall was excavated in 1959 [36]. Along with the rest of the cellar complex, the investigated hall probably played only a secondary functional role and was not a representative space of any kind, as Figure 2 shows through the absence of decorative architectural treatment, robustness of the construction, tolerance for high humidity, and aggressive environmental conditions. Nevertheless, it was located right beneath Diocletian’s residential complex in the southern quarter, showing the splendour—and the acoustic potential—of the once-existing emperor’s residence above [36].
Figure 2.
Photos of the interior of the single-nave hall; the apse in the north wall (left) and the rear of the hall (south wall), with the dodecahedral omnidirectional sound source used in measurements shown in the foreground (right).
The walls of the hall are made of massive, high-quality limestone blocks quarried on the island of Brač and were constructed using the opus quadratum technique. The upper zone was built using the opus mixtum technique, and the barrel vault ceiling is semi-covered in tufa. The hall is untypically small compared to Roman venues, most likely appropriate for a castrum-sized palace/settlement. As shown in Figure 3, it is 20.86 m long and 6.99 m wide, with the apex of the barrel vault reaching 6.96 m in height. An apse with a curvature radius of 2.22 m is built into the northern wall, with its highest point at 6.16 m above the floor. All four sides of the hall have openings in both the lower and upper zones that lead to adjacent spaces. The total volume of the hall is around 950 m3.
Figure 3.
Drawings of a single-nave hall with an apse (based on the drawings by Miljenko Žabčić, M. Geod. [39]); all dimensions are given in centimetres (in accordance with relevant building codes in Croatia).
2.2. Overview of Research Stages
The characterisation of the hall as a part of the cellar structure beneath Diocletian’s palace in terms of room acoustics was conducted in three stages. Initially, acoustic measurements were performed in the unoccupied and empty hall in terms of obtaining a set of monaural impulse responses and the values of standard reverberation and clarity parameters. In the second stage, an acoustic model of the hall was developed and calibrated using the results of in situ measurements. The third stage was orientated towards analysing the acoustic suitability of the hall as a potential venue for speech- and/or music-based events without the aid of a sound reinforcement system. The acoustic quality of the hall was investigated and evaluated using contemporary standardised metrics in terms of room acoustics [16] and speech intelligibility [27]. The results were discussed and compared with the ones obtained from similar archaeoacoustic studies of heritage sites.
2.3. Measurements
The measurements were carried out following the provisions of the ISO 3382-1:2009 standard [16] using the integrated impulse response method, leading to a set of impulse responses as the raw measurement data, obtained for all the defined source-receiver combinations. An omnidirectional dodecahedral loudspeaker, designed and produced by the research group, that meets the requirements on directivity laid out in ISO 3382-1 [16] was used as the sound source. An exponential sine sweep signal was used as the excitation. The excitation signal was limited to the frequency range from 45 Hz to 20 kHz to help protect the loudspeaker from operating at very low frequencies.
The core component of the measurement system was a Lenovo ThinkPad T14s laptop computer (Lenovo, Beijing, China) with a Focusrite Scarlett 2i2 third generation external audio interface (Focusrite Audio Engineering Ltd., High Wycombe, UK) connected to it. ARTA v1.9.7 measurement software [40] was installed on the computer and used for impulse response measurements. A class-D Crown XLS1500 power amplifier (Crown Audio, Northridge, CA, USA) was connected to the output of the audio interface and used to provide the amplified excitation signal to the sound source. The response of the hall was captured by a ½-inch class-1 GRAS 40AE omnidirectional measurement microphone connected to a GRAS Type 26CA preamplifier (GRAS Sound & Vibration, Holte, Denmark) and then to the input of the audio interface.
Figure 4 shows the positions of the sources (S1, S2, and S3), receivers (R1 to R12), and audience areas used in different stages of the study.
Figure 4.
Sound source positions (S1 to S3), receiver positions (R1 to R12), and audience areas used in different stages of the study; all dimensions are given in centimetres (in accordance with relevant building codes in Croatia).
Two positions were chosen for the sound source. The primary position S1 was set 2 m in front of the front wall of the hall with a half-cylindrical apse as a position that is likely to be occupied by a speaker addressing the audience or a performer involved in a musical performance. The secondary position S2 was chosen inside the apse itself, on the longitudinal axis of symmetry common to the hall and the apse, 1 m from the apse wall. The height of the acoustic centre of the sound source was set to 1.70 m in both positions. These source positions were used in measurements and model calibration. Source S3 shares the position of the source S1, but it is a directional source used in the analysis stage for emulating a live speaker.
A total of 12 microphone positions were chosen in the hall to cover the floor area, as displayed in Figure 4. The measurement microphone was mounted on a stand at the height of 1.25 m, i.e., the average ear height of people in a seated audience.
As described in detail further in the text, two different audience areas were defined: one for the hall configured for speech-based events and the other for music performances. These audience areas were used in the acoustic analysis of the hall in terms of its suitability for the said types of events.
The measurements were made in a sequence by placing the source in position S1 and moving the microphone through receiver positions until all the defined positions had been covered. After that, the source was relocated to position S2, and the measurement sequence was repeated.
The limiting factor during measurements was the amount of background noise. The hall is a part of the Diocletian cellars complex as one of the locations of the City Museum of Split. Thus, access to the hall was granted only during office hours. With visitors present, and the impossibility of separating the hall from the rest of the complex, noise coming from other spaces in the complex was constantly present in the hall during measurements. Moreover, visitors could freely enter the hall to have a look and/or pass through on the way to other spaces in the complex. This disturbance caused delays in the measurements, as the researchers were forced to wait for “silent” periods with less noise, repeat most of the measurements, and cease with measurement activities altogether when visitors entered the hall.
As the hall is located underground, the meteorological conditions during measurements were stable. In terms of indoor temperature and relative humidity, the former was found to be 17 ± 1 °C, and the latter was 49 ± 2% during the entire measurement period.
The measurements resulted in 24 room impulse responses for all combinations of two sound source positions and twelve receiver positions. These impulse responses were exported from ARTA as 32-bit floating-point wave audio files and loaded as raw data input into ODEON 18.16 Auditorium [41], installed on the same Lenovo T14s laptop computer that was used in the measurements. The computer runs on the Windows 11 Education v25H2 operating system (Microsoft Corporation, Redmond, WA, USA) and is equipped with an AMD Ryzen 7 PRO 5850U processor with Radeon Graphics (Advanced Micro Devices, Inc., Santa Clara, CA, USA) and 32 GB of RAM.
In ODEON [41], the frequency-dependent values were obtained for the standard set of room acoustic parameters associated with reverberation and clarity, namely the reverberation time T15 and T20, the early decay time EDT, speech clarity C50 and music clarity C80, and centre time TS.
Due to the described constraints imposed by noise, the reverberation time T30 was ultimately not calculated and displayed, as its values could not be obtained from most of the measured impulse responses due to insufficient dynamic range. The frequency-dependent reverberation time T20, centre time TS, and speech clarity C50 served as the main control parameters in the calibration of the hall model. Their measured values, averaged over all source-receiver combinations, are displayed further down in Section 2.4.
2.4. Development and Calibration of the Acoustic Model of the Hall
The three-dimensional geometrical model of the hall was made in SketchUp 2021, based on available floor plans and cross-sections obtained from [39] in AutoCAD2024 digital format and further verified by in situ measurements. The model was made with a 1 cm precision, as standard in civil engineering in Croatia. The surfaces in the model were assigned to different layers to facilitate easy assignment of materials to them in the modelling stage. As the hall is physically connected and acoustically coupled to adjacent spaces via passageways and openings in the walls, coarse representations of these spaces were added to the model, as shown in Figure 5. Irregular shapes such as the barrel vault, the apse wall, and the apse half-dome were simplified and represented with a small but sufficient number of three- and four-sided faces.
Figure 5.
Geometric 3D SketchUp model of the present state of the hall and the coupled adjacent spaces within the coherent cellar complex, presented in perspective view from the northwest.
The geometric model of the hall and the spaces adjacent and directly coupled to it, as shown in Figure 5, was prepared for importing into ODEON 18.16 Auditorium by means of the SU2ODEON v3.01 export plugin and was used as the basis for the development of the virtual acoustic model and its calibration [41].
As the entire complex was essentially buried for centuries and used as a sewage system in mediaeval times, it is impossible to assess how much damage such practice had caused and what the original appearance of the hall might have been, especially since there are no available historical records on the original design. Therefore, the development and calibration of the model were focused on recreating the present state of the hall. All surfaces were initially assigned a corresponding material with acoustic properties described by the values of sound absorption coefficient and scattering coefficient. In accordance with the physical structure of the hall, four different materials were defined. The floor of the hall has been resurfaced with smooth concrete, which also houses and/or covers electrical installations and drainage systems. The walls of the hall were made of massive stone blocks carved to size and shape and laid on top of each other with no mortar, using the opus quadratum technique, and were sourced from high-quality, dense limestone with low porosity [34]. The wall surfaces are neither perfectly flat nor smooth, as the joints between blocks are visible, and the exposed surfaces of the blocks have a rough structure due to the hand carving process. Passageways were left in the walls to connect the hall with adjacent spaces. The upper structure of the hall that rests on stone walls includes the barrel vault above the hall, the half-dome above the apse, and the openings that connect the hall to adjacent spaces. Two different materials were defined for the upper structure. One of them was assigned to the half-dome, the openings, and the perimeter of the barrel vault that rests directly on the walls. All these structures were built of mortar, stone and bricks using the opus mixtum construction technique. The resulting surfaces are considerably more rugged and irregular than the surface of the walls, thus justifying the assignment of higher absorption and scattering coefficients to this material. The rest of the barrel vault is covered with tufa. The surface is highly irregular, coarse, and full of cavities of various sizes, resulting in considerable absorption and scattering being assigned to the vault surface covered with tufa.
In accordance with the data on sound absorption of stone and other heavy masonry materials and structures found in the ODEON material database [42], as well as in literature [43,44], an assumption was made that the sound absorption coefficient of all four materials monotonically increases with frequency. The initial sound absorption coefficient of the concrete floor was adopted directly from ODEON. The stone walls were assigned the sound absorption coefficient in the range from 0.04 at 125 Hz to 0.10 at 8000 Hz. The opus mixtum surfaces were given the sound absorption coefficient of 0.10 at 125 Hz to 0.20 at 8000 Hz. Finally, a sound absorption coefficient of 0.10 at 125 Hz to 0.30 at 8000 Hz was assigned to the tufa covering the vault of the hall. The scattering coefficient was assigned according to recommendations given in the ODEON manual [42]: 0.02 for the concrete floor, 0.07 for the stone walls, 0.15 for the opus mixtum structure, and 0.30 for the tufa vault. The same absorption and scattering properties were assigned to the materials in the simplified representation of adjacent spaces attached to the hall.
The calibration of the model was made based on measured impulse responses for all source-receiver combinations, following the calibration methodology established in [45]. Two source positions and twelve receiver positions were defined in the model to match the ones used in measurements. The defined source and receiver positions are shown in Figure 4.
The impulse responses obtained for the 24 source-receiver combinations were loaded into ODEON. Frequency-dependent values of the early decay time EDT, reverberation time T15 and T20, centre time TS, and speech and music clarity C50 and C80, respectively, were calculated from these impulse responses and set as the target values for the optimisation/calibration procedure. The run-to-run variation between repeated acoustic simulations was examined on the initial model of the hall in terms of running the multipoint calculation in ODEON thirty times. The summary statistics of frequency-dependent reverberation and clarity parameters, both averaged over all positions and observed individually, have shown no run-to-run variation.
As the sound absorption properties for all materials except the concrete floor were only assumed but not exactly known, the initial calibration procedure was performed using the Genetic Material Optimizer as a feature of ODEON [41]. All the parameters listed above were included in the genetic optimisation process. The 63-Hz octave band was not considered due to high uncertainty observed both in measurements and in simulations based on geometrical acoustics, as well as to the lack of reliable data on low-frequency sound absorption.
The automatic calibration procedure provided useful insights but also violated the initial assumptions and yielded illogical and unsound results. In particular, higher absorption was assigned to stone walls than to the tufa ceiling, and the sound absorption coefficient of individual materials did not exhibit a monotonic increase but has shown either excessive low-frequency absorption or excessive mid-frequency absorption in a single octave band.
To address the issue, iterative manual calibration of the model was performed by adjusting the sound absorption properties of the stone, opus mixtum, and tufa. The same set of reverberation and clarity parameters was considered as in the genetic calibration stage. Emphasis was put on the reverberation time T20, the centre time TS as the general indicator of the temporal structure of the impulse response, and speech clarity C50 as the chosen clarity parameter, while monitoring all the remaining parameters.
The results of the calibration of the model are shown in terms of frequency-dependent values of the three chosen parameters, averaged over all source-receiver combinations, obtained from measurements and the calibrated model. The data is shown numerically in Table 1, Table 2 and Table 3. Apart from average values, standard deviations are displayed as well. The relative difference between calibrated and measured values, expressed as the percentage of “true” measured values, is shown for T20, whereas the absolute difference is shown for TS and C50. The obtained differences are also shown in terms of the corresponding JNDs, as defined in ISO 3382-1 [16].
Table 1.
Frequency-dependent reverberation time T20 averaged over all combinations of source and receiver positions (standard deviation shown in parentheses), obtained for the empty hall from (a) measurements and (b) simulation/calibration.
Table 2.
Frequency-dependent centre time Ts averaged over all combinations of source and receiver positions (standard deviation shown in parentheses), obtained for the empty hall from (a) measurements and (b) simulation/calibration.
Table 3.
Frequency-dependent speech clarity C50 averaged over all combinations of source and receiver positions (standard deviation shown in parentheses), obtained for the empty hall from (a) measurements and (b) simulation/calibration.
The results of the calibration procedure show that the best fit of the calculated values to the measured ones was obtained for the reverberation time T15, averaged over all receiver positions, as the difference was kept within one JND. The difference of centre time TS was kept around one JND, with systematic overestimation of its values in all frequency bands compared to the measured ones. The difference between the calculated and measured values of the early decay time EDT, the reverberation time T20, and the clarity parameters C50 and C80, averaged over all positions, was maintained within 1.5 JND.
The EDT was overestimated by calculations, whereas the T20, C50, and C80 were underestimated. For all parameters, exceptions from the stated findings were observed in individual receiver positions and frequency bands. The general limitation of the calibration procedure is observed in the octave band at 125 Hz, in which the reverberation parameters were overestimated in the simulations by 7 to 8 JND, and the clarity parameters were underestimated by 2 JND. As the only solution would have been to assign unrealistically high sound absorption in this frequency band to some or all the materials and surfaces, the decision was made to maintain the calibrated model as described, while acknowledging the limitations of the calibration procedure imposed by limitations of geometric acoustics modelling at low frequencies [46].
The sound absorption coefficient of both wooden chairs and the audience seated on them, as the foreseeable scenario in the occupied hall, was adopted directly from the ODEON material database [41], and the scattering coefficient of 0.7 for the audience area was set according to recommendations given in ODEON software [41,42]. The final absorption and scattering coefficients assigned to the materials and the corresponding surfaces in the model are shown in Table 4.
Table 4.
Sound absorption and sound scattering coefficients used in acoustic modelling.
2.5. Acoustic Analysis and Evaluation of the Hall
The developed model of the hall, calibrated to measured data, was used as the basis for an acoustic analysis of the hall in terms of its suitability as a venue for speech- and music-based events that are not aided by a sound reinforcement system but rely only on sound generated by a live source, i.e., a speaker or a musician. The analysis was based on calculations performed in ODEON 18.16 Auditorium [41].
Four different scenarios were investigated in terms of the type of event (speech- or music-based events) and occupancy (unoccupied or fully occupied hall). To account for the influence of occupancy, i.e., of the audience, two different audience areas were defined for speech- and music-based events, as shown in Figure 4. For speech-based events, the principal guideline for determining the number of seats in the audience was adopted from HRN DIN 18041:2024 [44], as the standard on acoustic quality of small and midsized rooms in terms of volume per seat. The stated literature stipulates the range of volume per seat of 6 to 8 m3/seat for mixed-purpose spaces (used for both speech and music). Considering the volume of the hall of 950 m3, its length and width, and the total floor area of 156 m2, the audience size that was deemed feasible was set to 135 seats, thus obtaining the volume per seat of 7.04 m3/seat. The audience area was divided into two sections, forming a total of 15 rows of seats with 9 standard-sized seats per row. For music-based events, the decision was made to reduce the audience area, as preliminary calculations revealed that the hall would otherwise not be reverberant enough for music performances. Given the requirement to not reduce the audience area to an unrealistically low number of seats that would not be appropriate for an audience attending a music performance, the audience area was ultimately reduced to 81 seats, thus obtaining the volume per seat of 11.72 m3/seat. The obtained value finds itself near the upper limit of the range defined in [44] for spaces for music, set from 7 to 12 m3/seat. In this case, the audience area was defined as a single section consisting of 9 rows of seats with 9 seats per row.
For this analysis, two extreme states of occupancy were examined. The hall in its unoccupied state was considered fully furnished with the required number of wooden chairs that matched the designated number of visitors. As the chairs introduce additional absorption into the hall, the resulting reverberation time in the unoccupied hall is lower than measured in the empty hall and obtained through calibration of the acoustic model of the hall. In the occupied state, all seats in the audience are taken, leading to even more absorption being added into the hall, resulting in a considerable reduction of reverberation time.
Following the intent to evaluate the acoustic quality of the hall in terms of contemporary criteria, the fundamental criteria regarding the optimal reverberation time based on room volume were derived from [44], nominally related to the occupancy of 80%. For speech-based events, the assumption was made that one-way communication would be predominant, in terms of a speaker addressing the audience, leading to the categorisation of the hall as a type-A2 space according to [44] and to the optimal reverberation time of 0.96 s. For music-based events, the hall was categorised as a type-A1 space with an optimal reverberation time of 1.41 s [44].
The settings in ODEON related to model parameters and calculations were adopted from the calibration stage and kept constant, thus facilitating a direct comparison of the obtained results. Since the measured reverberation time is found below 2 s, the target length of the calculated impulse response was set to 3500 ms. As the model includes not only the hall, but also adjacent spaces coupled with it, the number of late rays was considerably increased compared to the recommended value and was set to 250,000. All other options in the ODEON “Room Setup” menu were set to default settings.
The calculations were divided into two fundamental parts. The first part was aimed at determining the acoustic suitability for speech, and the second one was orientated towards determining the same for music performances.
2.5.1. Acoustic Suitability for Speech
Acoustic conditions in the hall in terms of suitability for speech-based events were examined in terms of (1) reverberance represented with reverberation time T20, (2) clarity represented with speech clarity C50, (3) speech intelligibility represented with speech transmission index STI, and (4) audibility and level uniformity represented with the overall A-weighted level SPL(A).
The source-receiver setup that was used both in measurements and in modelling of the empty hall was retained in this set of calculations. Omnidirectional source S1 was used only for calculating the reverberation time, while a directional source of speech S3, matching the S1 position, was used in all other calculations. Source S3 was defined to represent a speaker speaking in a raised voice and was orientated towards the audience. The directional and spectral characteristics of this source were taken from the ODEON database of sound sources [41].
Multipoint calculations performed for the 12 receiver points were used for obtaining the spatially averaged frequency-dependent values of relevant parameters, namely T20 and C50.
Grid calculations were used for obtaining the spatial distributions of single-number values of relevant parameters across the defined audience areas, namely C50,mid, and SPL(A) calculated according to [16] and STI calculated according to IEC 60268-16:2020 [27]. Grid height was set to 1.25 m above the floor to match the ear height of a seated listener. The spacing between receivers in the grid was set to 0.90 m, resulting in a total of 75 receivers in the two sections of the audience area defined for speech.
To quantify and interpret the results of grid calculations, corresponding summary statistics were provided for all the calculated parameters in terms of their minimum, maximum, and median values, along with selected percentile values, all stemming from respective cumulative distribution functions.
All parameters were calculated for the unoccupied and fully occupied hall.
The calculations of the STI were made for the no-noise scenario as the best-case scenario.
2.5.2. Acoustic Suitability for Music Performances
Acoustic conditions in the hall in terms of suitability for music performance were examined in terms of (1) reverberance represented by reverberation time T20, (2) clarity represented by music clarity C80, (3) sound strength and loudness represented by sound strength G, and (4) lateral support to perceived source width represented by early lateral energy fraction JLF.
The source-receiver setup that was used both in measurements and in modelling of the empty hall was retained in this set of calculations. Omnidirectional source S1 was used for all calculations, as required by [16]. Due to a large number of potential sources of musical content, no specific directional source was used.
As above, multipoint calculations performed for the 12 receiver points were again used for obtaining the spatially averaged frequency-dependent values of relevant parameters, in this case T20, C50, and G.
The spatial distributions of single-number values of relevant parameters across the defined audience areas were obtained from grid calculations. The parameters of interest were C50,mid, Gmid, and JLF,average, all calculated according to [16]. Grid height was left at the height of 1.25 m above the floor, and the spacing between receivers in the grid of 0.90 m was also retained from previous calculations, thus yielding a total of 40 receivers in the single-section audience area defined for music performances.
As defined in the previous section, summary statistics were provided for all grid calculations.
All parameters were calculated for the unoccupied and fully occupied hall.
3. Results
The hall under investigation was analysed from an acoustic perspective in terms of its suitability for two fundamentally different types of acoustic content, namely (1) speech and (2) music performance, both delivered naturally, i.e., without the aid of a sound reinforcement system. In each of the two cases, analysis has been made in terms of displaying and discussing the obtained values of acoustic parameters relevant for characterisation of acoustic properties of closed spaces. The influence of occupancy was examined in extremes, i.e., acoustic properties were obtained for (1) a fully furnished but unoccupied hall and (2) a fully occupied hall with all the seats in the audience taken.
3.1. Acoustic Parameters Relevant for Speech
The suitability of the hall for speech and speech-based events was analysed according to four criteria: (1) reverberation as the general indicator of acoustic conditions in a room, (2) speech clarity, (3) speech intelligibility, and (4) audibility of speech and the uniformity of its level in the audience, if produced by a live speaker relying only on their voice.
Due to the location and configuration of the hall, an assumption was made that the hall is to be used for events that employ one-way communication in terms of a speaker addressing the audience.
The reverberation in the hall was examined in terms of reverberation time T20. As stated in Section 2.3, optimal reverberation time for speech of 0.96 s was determined for the hall based on its volume and intended purpose, in accordance with the relevant standard used in Croatia for design and evaluation of acoustics of small and mid-sized rooms [44]. The size of the audience area in terms of the number of seats, in this case set to 135, was determined according to the same standard based on the size/volume of the hall.
The spatially averaged frequency-dependent values of reverberation time T20 obtained for the unoccupied and fully occupied hall are shown in Figure 6, plotted against the frequency-dependent tolerance mask defined and calculated based on the optimal reverberation time according to [44]. The values are also shown numerically, while the corresponding standard deviations are shown as error bars.
Figure 6.
Spatially averaged frequency-dependent reverberation time T20 (standard deviations shown as error bars), calculated for the unoccupied and occupied hall with the audience area set up for speech-based events, plotted against the tolerance mask defined according to HRN DIN 18041:2024 [44].
The single-number mid-frequency reverberation time T20,mid was calculated according to [16] and was found to be 0.97 s for the occupied hall. In the unoccupied hall, the same single-number quantity takes the value of 1.39 s.
To examine the uniformity of reverberation time across the frequency range of interest, bass ratio BR and treble ratio TR were calculated as well. The bass ratio was found to be 1.33 in the unoccupied hall and 1.61 in the fully occupied hall, whereas the treble ratio was found to be 0.66 and 0.73, respectively.
The clarity observed in the hall was calculated in terms of speech clarity C50 as the indicator of the energy ratio of the early and late parts of the impulse response relevant for speech, for the purpose of investigating the interaction between the speaker and the hall itself. For this reason, a sound source with directional and spectral characteristics representative of a speaker addressing the audience in a raised voice was used for obtaining the results presented below.
The spatially averaged frequency-dependent values of speech clarity C50 obtained for the unoccupied and fully occupied hall from the multipoint response calculations are shown on the chart in Figure 7. Mean values are also shown numerically, and the corresponding standard deviations are displayed using error bars. The average single-number mid-frequency speech clarity C50,mid was calculated according to [16] and was found to be −2.5 dB in the unoccupied hall and 1.1 dB in the occupied hall.
Figure 7.
Spatially averaged frequency-dependent speech clarity C50 (standard deviations shown as error bars), calculated for the unoccupied and occupied hall with the audience area set up for speech-based events.
The early part of the impulse response strongly depends on the distance between the source and the receiver, thus greatly affecting speech clarity. To illustrate this dependence, spatial distribution of the single-number mid-frequency speech clarity C50,mid was calculated according to [16] over the entire audience area. The results of these calculations are shown in Figure 8. In addition, Table 5 shows the accompanying summary statistics of C50,mid to help quantify the difference between the two cases of occupancy.
Figure 8.
Spatial distribution of single-number mid-frequency speech clarity C50,mid in the unoccupied (top) and occupied (bottom) hall set up for speech-based events.
Table 5.
Summary statistics of the spatial distribution of single-number mid-frequency speech clarity C50,mid in the unoccupied and occupied hall set up for speech-based events.
The speech intelligibility that can be achieved in the investigated configuration of the hall was examined in terms of calculating the speech transmission index STI according to the provisions of the relevant standard IEC 60268-16:2020 [27]. A situation in which a speaker addresses an audience without the aid of a sound reinforcement system was simulated, as it would have been the only option in historical terms, but also quite possible in present times. For this reason, the same sound source as the one used for speech clarity calculations was used here as well, as a representation of a speaker addressing the audience in a raised voice.
Due to the hall being a part of a larger complex as a historical attraction constantly visited by tourist groups and individual visitors, the noise profile in the hall is changing rapidly and frequently, making it impossible to determine the representative level and spectral content of background noise in the hall. Therefore, all calculations of the STI were performed for the no-noise scenario as the most favourable case, i.e., minimal background noise is assumed.
To illustrate how speech intelligibility depends on position in the audience area, spatial distribution of the STI over the entire audience area is shown in Figure 9 for the unoccupied and the fully occupied hall. The ranges of the STI that correspond to five established intelligibility ratings, as indicated on the colour bar in Figure 9, are defined in IEC 60268-16:2020 [27] for native listeners with normal hearing. The summary statistics of the STI that corresponds to these two cases are presented in Table 6.
Figure 9.
Spatial distribution of the speech transmission index STI in the unoccupied (top) and occupied (bottom) hall set up for speech-based events.
Table 6.
Summary statistics of the spatial distribution of the speech transmission index STI in the unoccupied and occupied hall set up for speech-based events.
The audibility of speech in terms of its overall level produced by a realistic speaker in the audience area and its uniformity in terms of the consistency of the level in the audience area were examined in the investigated configuration of the hall in terms of calculating the overall A-weighted sound pressure level produced by that speaker relying only on their own vocal effort. The same sound source was used as the one used for speech clarity and speech intelligibility calculations, representing a speaker speaking to the audience in a raised voice.
To illustrate how the loudness of the spoken word expressed in terms of the A-weighted sound pressure level depends on the position of the listener in the audience area, spatial distribution of the said parameter over the entire audience area is shown in Figure 10 for the unoccupied and the fully occupied hall. The accompanying summary statistics of the investigated parameter are presented in Table 7.
Figure 10.
Spatial distribution of the overall A-weighted sound pressure level in the unoccupied (top) and occupied (bottom) hall set up for speech-based events.
Table 7.
Summary statistics of the spatial distribution of the overall A-weighted sound pressure level in the unoccupied and occupied hall set up for speech-based events.
3.2. Acoustic Parameters Relevant for Music Performances
The suitability of the hall for music performances was analysed according to four criteria: (1) reverberation as the general indicator of acoustic conditions in the hall, (2) sound strength/gain provided by the hall and the resulting loudness of different vocal and instrumental sources, (3) music clarity, and (4) early lateral energy fraction.
The reverberation in the hall was examined in terms of reverberation time T20 by applying the same procedure as was performed for the assessment of the suitability of the hall for speech.
As described in Section 2.3, optimal reverberation time for music of 1.41 s was determined for the hall based on its volume and intended purpose, according to the provisions of the relevant standard [44]. The results of acoustic calculations made for the hall set up for speech-based events revealed that the hall would not be reverberant enough for music performances if the size of the audience area defined for speech was retained. To address the issue, a larger volume per person was allowed, according to [44], thus reducing the audience area to a realistic and still feasible size of 81 seats total. The audience area was set up as a single section and was set in the central part of the hall, away from the source, so that the audience would enjoy a more balanced blend of direct and reverberant sound.
Figure 11 displays the spatially averaged frequency-dependent values of reverberation time T20 calculated for the furnished but unoccupied hall and for the fully occupied hall and plotted against the tolerance mask defined according to [44], with standard deviations shown as error bars.
Figure 11.
Frequency-dependent spatially averaged reverberation time T20 (standard deviations shown as error bars) obtained for the hall set up for music performances, plotted against the tolerance mask defined according to HRN DIN 18041:2024 [44].
The single-number mid-frequency reverberation time T20,mid of 1.16 s was calculated for the occupied hall, whereas its value in the unoccupied hall reaches 1.53 s.
The uniformity of reverberation time in the frequency range of interest was investigated once again by calculating the bass ratio BR and the treble ratio TR. In the unoccupied hall, the bass ratio was found to be 1.33, and the treble ratio of 0.64 was found, while in the occupied hall the values changed to 1.55 and 0.70, respectively.
To investigate the support the hall provides to a performer in terms of sheer loudness, sound strength G was calculated as the indicator of the ratio of the energy of measured individual impulse responses in the hall to the ones that would have been measured in free field at a distance of 10 metres using the same sound source, according to ISO 3382-1 [16]. To obtain a generalised value of sound strength in the hall, while acknowledging that many different sound sources can take part in a musical performance, an omnidirectional source was used for calculations of sound strength, in accordance with [16].
Figure 12 presents the frequency-dependent values of sound strength G averaged over all 12 receiver positions used in the multipoint calculations. The results are shown for the two extreme cases of occupancy. As established in the presentation of other parameters, average values are complemented with the corresponding standard deviations displayed as error bars.
Figure 12.
Frequency-dependent spatially averaged sound strength G (standard deviations shown as error bars) calculated for the hall set up for music performances.
As a single-number value of sound strength for the entire audience area, to be used in subsequent loudness estimates, Gmid was found to be 15.5 dB for the empty hall and 13.5 dB for the fully occupied hall.
The changes of sound strength within the audience area were examined using the spatial distribution of the single-number mid-frequency strength Gmid over the defined audience area. The distribution of the values obtained from grid calculations in ODEON is shown in Figure 13. The corresponding summary statistics of Gmid are shown in Table 8 for the unoccupied and fully occupied hall.
Figure 13.
Spatial distribution of single-number mid-frequency sound strength Gmid in the unoccupied (top) and occupied (bottom) hall set up for music performances.
Table 8.
Summary statistics of the spatial distribution of the single-number mid-frequency sound strength Gmid in the unoccupied and occupied hall set up for music performances.
In terms of clarity of the performed musical pieces, music clarity C80 was calculated as the indicator of the energy ratio of the early and late parts of the impulse response with the boundary between the two parts set to 80 ms after direct sound. As stated above, the actual sound source in music performances can, in theory, be any type of vocalist or musical instrument or a group of instruments with different directional and spectral characteristics representative for each individual source. To overcome this issue and still gain an understanding of how music clarity changes in the hall with frequency and distance from the source, an omnidirectional source was used for calculating music clarity, while acknowledging that considerably higher clarity than calculated can be achieved in the audience with a highly directional sound source, e.g., a trumpet facing the audience.
The frequency-dependent values of music clarity C80, averaged over all 12 receiver positions, are shown in Figure 14 for the unoccupied and fully occupied hall. The standard deviations in frequency bands are shown as error bars. The average single-number mid-frequency music clarity C80,mid was calculated according to [16] and was found to be 0.6 dB in the unoccupied hall and 3.0 dB in the occupied hall.
Figure 14.
Spatially averaged frequency-dependent music clarity C80 (standard deviations shown as error bars) in the hall set up for music performances.
The dependence of music clarity on the distance from the source is illustrated with the spatial distribution of the single-number mid-frequency music clarity C80,mid over the audience area. This parameter was calculated according to [16]. The resulting distributions of the values for the unoccupied and fully occupied hall are shown in Figure 15, and the accompanying summary statistics of C80,mid are shown in Table 9.
Figure 15.
Spatial distribution of single-number mid-frequency music clarity C80,mid in the unoccupied (top) and occupied (bottom) hall set up for music performances.
Table 9.
Summary statistics of the spatial distribution of the single-number mid-frequency music clarity C80,mid in the unoccupied and occupied hall set up for music performances.
To quantify the support provided by early lateral reflections to the audience, thus evoking the perception of an apparently wider sound source, early lateral energy fraction JLF was calculated as defined in [16], with an omnidirectional sound source as the excitation of the room, as required.
The dependence of early lateral energy fraction on the position in the audience is illustrated with the spatial distribution of its single-number value JLF,average over the audience area, calculated as the arithmetic mean of the values obtained in octave bands at 125 Hz to 1000 Hz in each grid point. The resulting distribution of the values is shown in Figure 16 for the unoccupied and the occupied hall, along with the accompanying summary statistics displayed in Table 10.
Figure 16.
Spatial distribution of single-number early lateral energy fraction JLF,average in the unoccupied (top) and occupied (bottom) hall set up for music performances.
Table 10.
Summary statistics of the spatial distribution of the single-number mid-frequency early lateral energy fraction JLF,average in the unoccupied and occupied hall set up for music performances.
4. Discussion
4.1. Overall Assessment of Suitability of the Hall for Speech
In view of reverberation time as the basic criterion for evaluation, the average mid-frequency reverberation time of 0.97 s obtained for the occupied hall almost perfectly matches the optimal reverberation time of 0.96 s for a type-A2 space of this size (volume) used for one-way speech communication, calculated according to [44]. As shown above, the reverberance of the hall would be excessively high for low or no occupancy conditions.
The frequency-dependent values of reverberation time reveal excessive low-frequency reverberance compared to the rest of the frequency range in the unoccupied state of the hall, and the issue is even more pronounced in the occupied hall, due to the audience contributing to the total absorption in the hall predominantly at middle and high frequencies (in octave bands at 500 Hz and above). The frequency profile of reverberation time in the occupied hall exceeds the upper tolerance limit in the low-frequency region (octave bands at 125 Hz and 250 Hz), while in the unoccupied hall this discrepancy is extended to the mid-frequency region (octave bands at 500 Hz and 1000 Hz) as well. By contemporary criteria defined in [44], the described imbalance would normally call for acoustic treatment measures focused on reducing low-frequency reverberance, thus obtaining a more balanced frequency dependence of reverberation time. However, the investigated hall is a part of Diocletian’s palace complex as a World Heritage site protected by UNESCO [35] and as such is regarded as heritage of the highest rank. In Croatia, such sites are highly protected by relevant legislation and are exempt from requirements imposed on buildings, including the ones regarding acoustics. In this light, no permanent acoustic treatment in the investigated hall would be permitted. However, portable acoustic elements would most likely be allowed as a viable solution.
Although not as pronounced, an imbalance between mid-frequency reverberance and high-frequency reverberance can be observed in the frequency profile of reverberation time. In the occupied hall the imbalance is mitigated by the audience that introduces the most additional absorption in the mid-frequency region.
The overall judgement is that the obtained frequency profile of reverberation time is not suitable for speech by contemporary criteria due to excessive low-frequency reverberation, although the mid-frequency single-number value obtained for the fully occupied hall indicates that the amount of reverberation is optimal for the hall of this size and purpose. While spaces intended for music performances are usually designed with moderately emphasised low-frequency reverberance associated with the perception of warmth, the same is tolerated in spaces used for speech, although balanced reverberance across the low- and mid-frequency range is preferred.
In terms of speech clarity, the introduction of the audience into the hall yields significant average improvement of speech clarity of approximately 3 dB. For the unoccupied hall, the averaged single-number value of C50,mid is −2.5 dB, and negative values of C50 are observed in all individual octave bands except the ones at 4000 Hz and 8000 Hz, indicating the dominance of reverberant sound over the direct sound that has a detrimental effect on speech intelligibility. In a fully occupied hall, the averaged single-number value of C50,mid is 1.1 dB, and negative values of C50 are now observed only at low frequencies, i.e., in octave bands at 125 Hz and 250 Hz. The introduction of the audience increases the total absorption in the hall, leading to the shortening of the reverberation tail and the decrease of the energy contained in the late part of the impulse response. As the highest quantity of additional absorption is introduced in the mid-frequency range, the improvement of C50 is the greatest in this frequency range and reaches 3.5 dB. The increase at low frequencies does not exceed 2 dB, and the improvement at high frequencies is reflected in the 2.5-dB increase on average.
The spatial distribution of C50,mid shows that the frontal part of the audience area will benefit from higher speech clarity than the rear part due to its proximity to the sound source, regardless of the level of occupancy. However, in the frontal audience area being close to the source, the change of C50 exceeds one JND, depending on the distance from the source, whereas in the rear area its values are essentially constant, i.e., the variation is maintained within one JND, as confirmed by the summary statistics shown in Table 5. The summary statistics also show that the difference in mid-frequency speech clarity in the unoccupied and the occupied hall is statistically significant by any criteria, as there is no overlap between the two corresponding distributions of C50,mid. The improvement of speech clarity is in the range of 4 JND, as defined in [16], and is, therefore, clearly perceivable.
Regarding the physical characteristics of the hall itself and their contribution to speech clarity, the hall is, in fact, rather small, quite narrow (only 7 metres wide), barrel-vaulted, and constructed of (at least nominally) acoustically reflective materials. All these characteristics should normally lead to strong early reflections as an aid to direct sound that contributes to enhanced speech clarity [47]. However, as the hall is elongated and has parallel lateral walls, the sound coming from the source positioned in front of the apse will be strongly reflected only from the front part of these walls close to the source itself, thus benefiting only the front audience area. For the rear audience area, the angle of incidence of sound is large (approaching grazing incidence), thus causing the early reflections to miss most of the rear audience area as defined in the model. In addition, the presence of openings and passages to adjacent spaces in the lateral walls causes a part of these early reflections to be lost as they enter these openings and propagate to adjacent halls. As for the ceiling, its barrel-vaulted shape would nominally be ideal for concentrating the early reflections in the audience area [47]. However, as shown in Figure 2, the ceiling follows Vitruvius’s stucco treatment guideline [5] but also has suffered considerable damage over the centuries, especially its top part covered with tufa, thus having considerable absorption and diffusion. As such, it cannot provide strong early reflections in the audience. The described physical properties of the hall ultimately lead to limited early energy in the audience, and, consequently, lower speech clarity.
Despite the barely acceptable values of speech clarity for a space that is to be used for speech-based events, a basic analysis of speech intelligibility in the unoccupied hall has shown that fair speech intelligibility (STI > 0.45) can be achieved in the entire audience area for native listeners with normal hearing even in the unoccupied hall, while a small part of the audience area closest to the speaker will experience good speech intelligibility (STI > 0.60). In a fully occupied hall, good speech intelligibility can be achieved in the entire audience area. A more detailed analysis is based on the nominal qualification bands defined for the STI in IEC 60268-16:2020 [27] that transcend the five basic intelligibility categories. According to this qualification, the range of STI from 0.64 to 0.72 achieved in the occupied hall corresponds to B (0.68 ≤ STI < 0.72) and C (0.64 ≤ STI < 0.68) qualification bands. Both bands designate high speech intelligibility appropriate for theatres, speech auditoria, parliaments, courts, and other venues where complex messages that contain unfamiliar words need to be verbally conveyed to the audience. The range of STI from 0.55 to 0.63 achieved in the unoccupied hall falls mostly within the E (0.56 ≤ STI < 0.60) and F (0.52 ≤ STI < 0.56) qualification bands and only marginally in the D (0.60 ≤ STI < 0.64) band. The D band designates good speech intelligibility appropriate for conveying complex messages using familiar words, e.g., in lecture halls, classrooms, or concert halls. The E and F bands designate speech intelligibility appropriate for conveying complex messages within a familiar context, with the E band being appropriate, e.g., for concert halls and modern churches, and the F band being sufficient for shopping malls, VA systems, and cathedrals [27].
The summary statistics of the STI show that the introduction of the audience into the hall leads to a statistically significant increase in speech intelligibility, as the distributions of the STI for the two examined occupancy conditions do not overlap. The improvement of speech intelligibility is considerable and is reflected in the increase of the STI by 0.09 on average, i.e., in the range of 3 JND, as determined in [48]. As analysed above, the introduction of the audience leads to improvement of speech intelligibility by two full STI qualification bands.
The obtained values of the STI are additionally analysed in terms of communication in a non-native language. As defined in IEC 60268-16:2020 [27], the thresholds of the STI for five basic levels of speech intelligibility are shifted upwards for non-native listeners, depending on their language proficiency level. In this context, experienced non-native listeners would still be able to perceive fair speech intelligibility anywhere in the audience area, regardless of occupancy, while good speech intelligibility would be achieved only in the few front rows closest to the speaker in the fully occupied hall. Non-native listeners with intermediate proficiency would experience fair speech intelligibility anywhere in the audience area in the fully occupied hall and only in 10% of the audience area closest to the speaker in the unoccupied hall. Good speech intelligibility would be impossible to achieve in this case.
In terms of audibility of speech, the plots displayed in Figure 10 and the summary statistics in Table 7 show that the overall A-weighted level of speech achieved in the audience area ranges from 57.9 to 62.4 dBA, indicating a high uniformity of level and, consequently, perceived loudness of the delivered spoken content. This observation is in line with the acoustic conditions in the hall dominated by reverberant sound due to the small amount of total sound absorption present in the unoccupied hall. The summary statistic for the unoccupied hall indicates that 50% of the audience area, i.e., its rear section, experiences a practically constant sound pressure level that stems from dominant reverberant sound, as the median value of SPL(A) is only 1.4 dB higher than the one calculated as the minimum. The other half of the audience area, i.e., the front section that is closer to the source, exhibits an increase in sound pressure level of 3 dB from the seats farthest from the source to the ones closest to it due to the increasing influence of direct sound from the sound source. Thus, in the unoccupied hall (or with low occupancy), a listener would be able to clearly hear the speech content everywhere in the audience, provided that the background noise is kept sufficiently low. However, for improved intelligibility, it would be preferable for them to sit in the front section of the audience.
Due to considerable additional absorption added to the hall by the presence of the audience, in a fully occupied hall the overall sound pressure level is reduced as a direct consequence of the reduction of the level of the reverberant component of sound. As shown in Figure 10 and in Table 7, the greatest reduction of approximately 4 dB is observed in the rear section of the audience located far away from the source, in which the overall level is dominated by reverberant sound. In the parts of the audience closest to the source, the observed reduction reaches only 2 dB due to the dominance of direct sound. The overall level now ranges from 53.7 to 60.2 dBA, indicating reduced uniformity of level compared to the unoccupied state. As before, the sound pressure level in the rear section of the audience, taking approximately 50% of the total audience area, is maintained within a small margin of 1.8 dB, whereas in the front section of the audience it increases by almost 5 dB as the distance between the listener and the source decreases.
In terms of both audibility of speech and uniformity of speech level, the absence of an audience (or in low occupancy conditions) inherently leads to better audibility and uniformity. On the other hand, the fully occupied hall provides more favourable acoustic conditions in terms of speech clarity and intelligibility. However, the reduction of the overall sound pressure level by 2 to 4 decibels is noticeable and easily perceivable as a reduction of the loudness of the delivered speech. Given that the audience itself can be an internal source of background noise, any reduction in speech level is detrimental to its audibility and impacts the ability of the audience to follow and comprehend the spoken content. As shown above, the rear section of the audience would suffer the greatest reduction of speech level in the occupied state of the hall, thus being more prone to the described difficulties.
The historical, architectural, and cultural significance of the hall and the entire complex of Diocletian’s palace provide a high potential for its revitalisation and possible use for a variety of events. In light of the four investigated contemporary criteria and the overall suitability of the hall for speech-based events, the overall judgement is that the hall in its present condition can be used for speech-based events that (1) predominantly employ one-way communication between the speaker and the audience, (2) assume high occupancy of the audience area, and (3) imply a low level of background noise in the hall, as well as a trained speaker capable of producing a sufficiently high speech level. As a few examples, the events befitting this hall could include poetry readings or literary recitals in general, as well as various sorts of presentations, meetings, openings of exhibitions, etc. However, for two-way verbal communication involving multiple speakers, e.g., mingling after the official part of an event has ended, the hall would need to undergo additional acoustic treatment to further reduce its broadband reverberance, on top of dealing with excessive low-frequency reverberance, as suggested above. Current conditions in the hall in terms of speech intelligibility would be suitable for events being held in a native language, whereas the intelligibility of speech delivered in a non-native language would be somewhat compromised. Given that the Diocletian’s palace complex and the city of Split in general receive a high number of foreign visitors each year, further improvement in speech intelligibility in the hall would facilitate the extension of the events happening in the hall to the foreign tourist population as well. As permanent acoustic treatment would be regarded as invasive in this kind of space, intelligibility could be improved by adding a sound reinforcement system composed of loudspeakers with higher directivity than the one exhibited by a live speaker. Such an addition would simultaneously (1) improve the direct-to-reverberant ratio, thus improving clarity and intelligibility, (2) improve speech-to-noise ratio in elevated background noise conditions, again leading to improved intelligibility, and (3) provide aid for untrained speakers with limited vocal ability (in terms of volume).
4.2. Overall Assessment of Suitability of the Hall for Music
In terms of reverberation, the results indicate that the reduction of the audience area from 135 seats for speech to 81 seats for music performances was insufficient to achieve optimal reverberance for music. In the unoccupied hall, the mid-frequency reverberation time exceeds the optimal value but is still within the defined tolerance. Thus, further reduction of the audience area would be needed to achieve a match of the calculated value and the one found to be optimal according to [44]. However, such a reduction would have to be significant, reducing the audience to 20 to 30 seats at most, which raises the question of the feasibility of music-based events, thus making such a step unlikely.
Viewed as a frequency-dependent parameter, the reverberation time again exhibits excessive low-frequency reverberance, leading to a perception of warmth provided by the hall in terms of acoustics, and the presence of the audience further emphasises the issue. Similar to the setup of the hall made for speech, the hall in the setup for music again displays lower high-frequency reverberance compared to the one observed in the mid-frequency region, thus leading to a reduced perception of brilliance and liveness and further adding to the perception of warmth. Figure 11 shows that the frequency profile of reverberation time suffers from a lack of uniformity. For the occupied hall, it is kept within the tolerance limits only in the three lowest octave bands (at 125 Hz, 250 Hz, and 500 Hz), whereas it falls below the tolerance zone in the rest of the frequency range. In the unoccupied hall the inconsistency of reverberation time is taken to the extreme, as it meets the tolerances only in the octave band at 1000 Hz. The values in the lower bands overshoot the tolerance range, while the ones in the higher bands fall below it.
Realistically, the perception of the acoustic situation in the hall would depend on the performer and the type of instrument they play. In terms of objective criteria, the mid-frequency reverberation time in the occupied hall is found near the lower limit of the range of typical values for performance spaces defined in [16] and is too low for music performances, according to [44]. In addition, the obtained frequency profile of reverberation time in the occupied hall is not suitable for music performances, as the hall is not reverberant enough in the mid- and high-frequency ranges and simultaneously displays excessive low-frequency reverberation. The solution to this issue would require (1) further reduction of the audience area to achieve sufficient mid-frequency and high-frequency reverberation, and (2) introduction of low-frequency absorption into the hall to reduce low-frequency reverberation and balance it out with reverberation in the rest of the frequency range. At present, neither step is viable, as (1) it would not be feasible to hold any kind of music performance for a very small audience, and (2) no permanent acoustic interventions would likely be allowed in the hall, as explained above, but only the ones of temporary nature.
In terms of sound strength calculated for the hall, the frequency dependence of sound strength reflects the frequency profile of reverberation time. The highest values of strength are found at low frequencies, both in the unoccupied and in the occupied hall, and the sound strength monotonically decreases with the increase of frequency. The introduction of the audience into the hall results in an average decrease in sound strength of 2 dB in the mid-frequency range, i.e., in the octave bands at 500 Hz, 1000 Hz, and 2000 Hz. According to [16], this difference is equal to 2 JND and is, therefore, perceivable. The reduction of sound strength in the rest of the frequency range is kept around 1 dB and is, therefore, just noticeable.
The spatial distributions of single-number sound strength Gmid show that the introduction of the audience into the hall results in a statistically significant decrease in sound strength, as the difference of the median values for the two cases of occupancy exceeds 3 dB, i.e., 3 JND. In terms of variation with the distance from the source, higher strength was observed in the frontal part of the audience area that is closer to the source, regardless of occupancy. In the rear half of the audience area, the values of Gmid are contained within 1 JND, suggesting that the entire rear half of the audience essentially experiences the same perception of loudness. A clear perceptual difference in terms of strength and loudness can be expected only for the 10% of the audience seated in the frontmost rows, compared to the rear half of the audience.
The suitability of rooms for different music performances and corresponding ensembles in terms of sheer loudness was assessed by considering the obtained values of strength Gmid for the entire audience area, the size of the room in terms of its net volume V, and reverberation in terms of reverberation time T. In the design stage of both performance [49,50] and rehearsal rooms [51,52], these quantities and the relationships between them are adjusted so that the room fits the performers (a soloist, a small ensemble, or a large ensemble) in terms of the level of the played sound. In the frame of this study, all these values are known for the investigated hall and are not subject to change, as no permanent acoustic intervention (or intervention of any kind) is allowed in the hall.
Given the volume of the hall of 950 m3, its floor area of 156 m2, and the size of the audience area of 81 seats, it is reasonable to assume that the hall would be suitable for music performed by a soloist or a small ensemble at most. According to the procedure described in the ISO 23591 standard [52], to achieve the target level at forte between 85 dB and 90 dB, and having in mind the average mid-frequency sound strength Gmid of 13.5 dB in the fully occupied hall, the sound power level of a soloist or an ensemble should be in the range from 102.5 dBre1pW to 107.5 dBre1pW, which translates to sound power in the range from 18 mW to 56 mW.
An example of a small ensemble that could be performing in the hall is the Dalmatian klapa, traditionally consisting of anywhere from 4 to 10 male singers, singing a capella, inscribed in 2012 on the Representative List of the Intangible Cultural Heritage of Humanity [53]. In recent decades, all-female and mixed klapas are present as well. Considering the average size of a klapa of 7 singers and taking the sound power data from [52], the total sound power of an average-sized klapa can be set to 25 mW, thus making the ensemble capable of reaching the forte level of 86.5 dB in the audience area.
A classic string quartet consisting of two violins, a viola, and a cello reaches the total sound power of 3.1 mW, thereby producing the level of only 77.4 dB in the audience area, which proves to be insufficient for the audience to perceive the loudness appropriate for playing at true forte but is more appropriate for an intimate performance.
A solo singer accompanied by a guitar is another traditional music form in Dalmatia. Together, these two sound sources yield the total sound power of 4.4 mW, which corresponds to the sound pressure level of 79 dB. However, the audience would likely expect a lower overall level from such a performer and a performance of such an intimate nature.
A solo trumpet has the sound power of 12.5 mW and, by the same calculations, would produce the sound pressure level in the audience of 83.5 dB, just shy of the lower limit set for optimal levels at forte. However, the calculations presented in [52] do not consider source directivity but are limited to the level of reverberant sound achievable in a room characterised by strength G and excited by an omnidirectional sound source with a known sound power. For highly directional sources such as the trumpet, the audience in this hall is expected to be in the direct sound field of the instrument, thus being exposed to direct sound, the level of which greatly exceeds the one calculated for reverberant sound.
Finally, even a small ensemble of three or four loud brass instruments would already exceed the sound power limit calculated above for this particular hall, proving to be excessively loud while playing at forte in this hall.
Similar to speech clarity calculated for the hall set up for speech-based events, a significant average improvement of mid-frequency music clarity of approximately 2.5 dB can be expected with the audience present in the hall. The frequency-dependent values of the spatially averaged music clarity shown in Figure 14 for the unoccupied hall reveal that negative values of C80 are limited to octave bands at 500 Hz and below, indicating the dominance of reverberant sound over the direct sound. With the hall fully occupied, a negative value of C80 is now observed only in the octave band at 125 Hz. Due to the sound absorption properties of the audience, the increase of C80 is again the greatest in the mid-frequency range, as stated above. The increase at low (125 Hz and 250 Hz) and high (4000 Hz and 8000 Hz) frequencies is kept below 2 dB. In terms of the typical range of music clarity defined in [16], i.e., from −5 dB to +5 dB for performance spaces, the median music clarity in the unoccupied hall is in the midpoint of the defined range at 0 dB, implying a balance between direct and reverberant sound. In the occupied hall, the same value reaches 2.5 dB, thus marking a shift towards clearer, but simultaneously drier sound.
Figure 15 shows the spatial distribution of the single-number mid-frequency music clarity C80,mid. Compared with clarity calculated for speech, some commonalities can be found. As found for speech, higher music clarity was found in the frontal part of the audience area, as it is closer to the source, both in the unoccupied and in the occupied hall. However, the reconfiguration of the audience area to a single section that is moved away from the sound source to the central part of the hall is reflected in the obtained values of C80,mid. While the range of C50,mid found in the audience area for speech was 4 dB in the unoccupied hall and 3.2 dB in the occupied hall, the C80,mid calculated for the entire one-section audience is kept within a 1.7-dB range in the unoccupied hall and within a 1.6-dB range in the occupied hall, just exceeding 1 JND in both cases. This suggests that the perceived clarity of music should be quite similar for all persons seated in the audience area. As with speech clarity, the difference in mid-frequency music clarity for the two extreme cases of occupancy is statistically significant by any criteria, as the summary statistics show no overlap between the two corresponding distributions of C80,mid. The improvement of music clarity is in the order of 2.5 JND, as defined in [16], thus making it clearly perceivable.
In terms of the early lateral energy fraction, its single-number value proves to be quite high. The median value of 0.290 was obtained for the unoccupied hall and 0.278 in the occupied hall, while individual values exceed 0.250 in the entire audience area. The typical range of values for performance spaces defined in [16] is from 0.100 to 0.350. In this context, the hall is expected to provide strong early lateral energy to the audience, thus contributing to the perceptive broadening of the sound source. High values of this parameter are attributed to the overall shape of the hall, as the ratio of length to width is approximately 3 to 1, and the width of the hall in absolute terms is only 7 metres. In addition, all the walls of the hall, including the lateral ones, are highly reflective as well, thus producing strong lateral reflections that arrive to the audience shortly after direct sound. The summary statistics reveal that the JLF,average is maintained within one JND of 0.05, as defined in [16], in the entire audience area, in both the unoccupied and the occupied hall. Moreover, the influence of occupancy on JLF,average is marginal, as the reduction caused by introducing the audience is far below the defined JND.
The spatial distributions of JLF,average shown in Figure 16 reveal interesting, localised effects of the layout of lateral walls on the JLF,average in the audience area, i.e., the front right part of the audience area displays higher values of JLF,average than observed in the front left part. The difference is minute and has been attributed to the position of the sound source and the audience area in the hall, as well as to the layout of the passages in both lateral walls towards adjacent spaces. These openings are essentially “black holes” for early reflections, as any sound that enters a passage will suffer multiple reflections from passage walls and ultimately travel into the adjacent space to which the passage leads. At first glance, the right wall has two passages in the critical area that provides lateral reflections for the audience, whereas the left wall has only one. Thus, the early lateral energy fraction was expected to be higher in the left part of the audience. However, additional analysis was performed in terms of calculating the spatial distributions of JLF,average for five alternative positions of the sound source chosen along the width of the hall. In addition, the paths of the first reflections from the lateral walls were traced from the source position to the audience area in each of the cases, including the original one. For the original source position, it was ultimately concluded that there are two clear paths for the first reflections from the lateral walls to the front right part of the audience, as the incident sound meets a flat surface on both lateral walls to be reflected from them into the audience. On the other hand, only one early reflection path has been identified for the front left part of the audience area, i.e., the first reflection arrives only from the left lateral wall, as the right lateral wall contains a passage exactly in the area sound would reflect from, thus eliminating this early reflection.
As discussed above, the value of the hall in terms of architecture and heritage and the fact that the entire complex of Diocletian’s palace is a unique heritage site should be the driving force for the use of the hall as a venue for music performances as well. Given that Dalmatia has a great tradition in group a cappella singing, a logical step would be to use the hall as a venue for klapa performances. The experience of listening to this kind of performance in a historically and architecturally unique space such as this one would undoubtedly be rewarding for a variety of audiences, both domestic and foreign. In the context of the investigated contemporary criteria and the overall assessment of the suitability of the hall for music performances, the hall in its current condition, with reduced audience area, is nominally not suitable for music performances in the occupied condition, as mid-frequency reverberation in the hall is shorter than required, while low-frequency reverberation is excessive. The gain provided by the hall and its overall size suggest that the hall can and should serve as a venue for solo or small-ensemble music performances. Positive values of music clarity in the occupied hall point to a clear sound provided by the hall in terms of the possibility for the listener to distinguish individual notes even in fast passages. Due to the size and shape of the hall, the size of the audience area, and the presumed types of music performances to take place inside, the obtained values of early lateral energy fraction are rather high. Small-hall performances lean towards clarity, intimacy, and communication between the performers and the audience, rather than an unnatural and overemphasized perceived width of the sound source. Given all the factors mentioned above, in this particular hall the perception of an excessively large sound source might disagree with its actual size and with the overall acoustic information received by a listener seated in the audience.
However, in the context of klapa singing, excessive low-frequency reverberance would provide a warmer tone to the entire ensemble by enhancing the baritone and, especially, bass voices that are often in a secondary role and lower in volume compared to the rest of the ensemble. As discussed above, the size of the hall and its reverberance yield a gain that is quite suitable for a small ensemble of singers such as a klapa. The clarity of sound in the hall would enable the listeners to clearly understand the lyrics of the song, as an important aspect of Dalmatian tradition, and perhaps even to sing along if appropriate. In terms of perceptual broadening of the sound source, in the case of klapa, it would not disagree with the actual size of the ensemble, as the singers are normally placed in a single row or arc and would occupy a substantial part of the actual width of the hall (depending on the number of singers).
4.3. Comparison with the Results of Similar Archaeoacoustic Studies
As the final step of the analysis, a comparison was made of the obtained results for the investigated hall with the results of similar archaeoacoustic studies. The greatest challenge in this analysis was to find a study performed on a heritage site similar to the one examined in this study. As the investigated hall is quite unique, this challenge has not been successfully overcome. Instead, a comparison of the investigated site was made with a wide range of heritage sites, ranging from prehistoric caves to ancient Greek and Roman theatres, catacombs, and early Christian churches.
As presented in previous sections, the investigated hall was evaluated regarding its potential use as a space for both speech and music. The values of the key parameters that were obtained for the speech configuration are: T20,mid = 1.39 s (unoccupied) and 0.97 s (occupied), C50,mid = −2.5 dB (unoccupied) and 1.1 dB (occupied), STI = 0.55–0.63 (unoccupied) and 0.64 to 0.72 (occupied), and SPL(A) = 58–62 dBA (unoccupied) and 54–60 dB (occupied). The values of the key parameters that were obtained for the music configuration are: T20,mid = 1.53 s (unoccupied) and 1.16 s (occupied), C80,mid = 0.6 dB (unoccupied) and 3.0 dB (occupied), Gmid = 15.5 (unoccupied) and 13.5 (occupied), and JLF,average = 0.277–0.317 (unoccupied) and 0.262–0.306 dB (occupied).
The type of heritage sites that are extensively analysed are the open-air theatres from the Greek and Roman eras. In [1], an extensive analysis was made in terms of the evolution of open-air theatres from the early era of Greek civilisation to Roman times, in terms of their construction and, consequently, acoustic properties. It was found that theatre construction has evolved with time in terms of increasing the level of speech in the audience area and lengthening the reverberation time. In terms of specific values, the obtainable speech level in the audience area has been increased to 53 dB at the point closest to the stage, i.e., 25 m away from the source. As the distance increases (to 65 m as the maximum), the level drops in a manner quite similar to the one observed in the free field. However, to achieve the stated levels, the sound pressure level of the source was set to “performance level”, i.e., considerably louder than speaking in a normal or even raised voice. The reverberation time was found in the range from 0.3 s for early designs to 1.5 s in Roman theatres. As for speech intelligibility, being an important feature for theatres, the speech transmission index STI was found to be in the range from 0.42 (poor) to as high as 0.80 (excellent). Specifically for Roman theatres, the values of the STI range from 0.45 to 0.53 in the unoccupied state and from 0.51 to 0.70 in the occupied setup, suggesting that the acoustics of Roman theatres is similar to the one aimed for in modern ones.
As the design of the investigated hall is fundamentally different from an open-air theatre, the comparison of the two designs is challenging, as their acoustic behaviour is quite different as well. As outdoor spaces, open-air theatres lack the late reflections and true reverberation tail in their impulse response, which is composed predominantly of direct sound and early reflections. Hence, they are characterised by a reasonably short reverberation time, high speech clarity and intelligibility, and non-uniform speech level distribution in the audience area. As a specific example, the Roman theatre of Gubbio investigated in [14] displays very high speech and music clarity in the fully open state, ranging from 13–14 dB and from 14–18 dB, respectively. Both parameters far exceed the values found in the investigated hall. On the other hand, the sound strength in the theatre is quite low (around 0 dB on average), as expected for a fully open space. An architectural intervention in the form of adding the upper porch and/or roof to partially close the otherwise open theatre led to a considerable reduction of speech clarity to the range from −2 to 3 dB and music clarity from 2 to 8 dB, thus being more in line with expected values for a space of this type. Simultaneously, the values of sound strength in the audience area increased to the range from 4 dB to 10 dB, thus contributing to the audibility of the spoken content. As another example of a Roman theatre investigated in [54], the theatre in Posillipo, Naples, displays the value of reverberation time T30,mid of approximately 1.3 s, while its STI values ranging from 0.58 to 0.75 are similar to those obtained in the investigated hall, as well as the values of music clarity of 1.5 dB to 5 dB. An investigation presented in [2] focuses on the Paphos theatre. The values of reverberation time of 1.75 s and RASTI ranging from 0.50 to 0.55 are again similar to those obtained in the investigated hall.
Enclosed spaces such as early Christian basilicas, catacombs, or even prehistoric caves are considerably more similar to the investigated hall in terms of the way the overall sound field forms inside the space. Yet, in terms of sheer size and shape, a suitable counterpart to the investigated hall has not been found in literature. In [2], a study of selected prehistoric caves has been made, and it was determined that the reverberation time inside is found in the range from 0.75 to 1.50 s, while the speech transmission index expressed as RASTI was found to be in the range from 0.62 to 0.77. As investigated in [31], the upper level of the Catacombs of San Gennaro, Naples, proved to be quite similar to the investigated hall in terms of shape (elongation) and stone cladding but considerably larger, of a more developed floor plan and including also porous earthen surfaces filled with burial niches. Nevertheless, the values of T30,mid of 1.14 s, and the STI of 0.63 mark this space as acoustically quite comparable to the investigated hall, along with the caves from the previous example. Similar values of reverberation time and speech transmission index in the catacomb and the investigated hall, despite a considerable size difference, point to a larger amount of total sound absorption in the catacomb attributed to a large number of earthen surfaces and the shaping of the walls through rows of burial niches.
Early Christian churches, researched in [6] in terms of acoustics and architectural typology, inherited the form of Roman basilicas, representing a potential reference for comparison with the investigated single-nave hall. An early domus with a small room volume of only 340 m3 was compared with larger basilicas (up to 75,000 m3), showing the values of T30,mid of around 1.0 s and STI 0.60, to be comparable to Roman and early Christian contexts, where speech intelligibility was often prioritised. Larger basilicas reflect later liturgical developments exhibited with long reverberation appropriate for ritual vocalisation of the liturgical chant [6,7,21], accompanied by poor speech intelligibility. As such, they are in no way comparable with the investigated hall or with other discussed types of heritage sites.
5. Summary and Conclusions
Archaeoacoustic examinations of heritage sites are carried out for a variety of reasons, ranging from acoustic reconstruction of sites that no longer exist to analysis of existing sites rated as top-level tangible heritage that could be used in a modern context considerably different from the original one. This study applied the latter approach by investigating a single-nave hall from the Roman era, located in the cellars of Diocletian’s palace in Split, Croatia, in terms of its acoustic properties and subsequently evaluating its potential for non-invasive future revitalisation in terms of contemporary use as a performance space. Detailed in situ acoustic measurements were performed in the hall, followed by the development and calibration of the virtual model of the hall. Based on the calibrated model, an analysis of the suitability of the hall for speech-based events and music performances was performed. As evaluation criteria, reverberation time, speech clarity, speech transmission index, and overall A-weighted sound pressure level were calculated for the speech-based configuration of the hall, while reverberation time, sound strength, music clarity, and early lateral energy fraction were calculated as criteria for music performances. The obtained values of listed parameters were discussed not only in objective terms but were also linked with contemporary perceptual dimensions such as speech intelligibility, audibility and level uniformity in the hall, as well as the perceived warmth and brilliance of its acoustic response.
The contribution of this research is not only in documenting the acoustic quality of the heritage site but also in further evaluation on how archaeoacoustics can support existing intangible heritage for continuous usage with minimal or no design interventions. In this study, the aim was not to propose acoustical treatment through reconstruction of the site but the revival of the intangible acoustic heritage of a preserved space of unknown function. The literature suggests that domestic architecture served as a space for a social performance and for presenting homeowners’ identities. While the function of the hall itself remains unknown, it may be presumed that the space was suitable for the period in which it was created, based on the performed acoustic analysis and the comparison with similar archaeoacoustic studies. Nevertheless, the contemporary value of the investigated hall lies in its ability to host various events due to its ability to adapt to either speech-based events, especially those based on one-way communication between a speaker and the audience, or smaller music ensembles, without compromising its character as a cultural heritage site worthy of UNESCO protection. The acoustics of the hall appears to be favourable for klapa singing as a UNESCO Intangible Cultural Heritage in itself, potentially leading to enhanced and even more compelling listener experience.
The limitations of this study are reflected in several open issues. Firstly, the cellars of the Diocletian’s palace have been excavated only recently, while the fieldwork and examination are still ongoing. Therefore, there is insufficient evidence to draw firm conclusions about the function of the hall as a part of this complex, its spatial layout, material treatment, and related aspects, including the uncertain preferences on acoustics valued by the Romans. Being a part of a unique complex with a dual function of a both a palace and a castrum, the hall is hardly comparable to other historic sites, as its exact function remains unknown, and its small volume and lack of ornaments is very atypical for Roman venues. Secondly, as the cellars form an important part of an ancient Roman and Croatian heritage site and are a part of the City Museum of Split, they are visited daily by many tourists, thus making it difficult to perform acoustic measurements due to persistent and constantly elevated background noise. Thirdly, the calculation methods implemented in the room acoustics simulation software are limited at lower frequencies that are crucial for the perceptual aspect of warmth. Finally, in the absence of other options, other acoustic research studies of historic sites have likewise been conducted based on modern standards which primarily reflect contemporary evaluation criteria, rather than those specific to the period in which the structures were built. The same limitations, in turn, indicate the need for further research.
The first two constraints can be additionally explored by extending the archaeoacoustic analysis to other parts of the cellar complex of the Diocletian’s palace to provide a comprehensive assessment of the acoustic quality of the entire complex. The remaining limitations, particularly those related to the analysis of acoustic parameters and the evaluation of acoustic quality itself, may be addressed through a human-centred approach. It would go beyond objective evaluation by employing binaural measurements, virtual acoustics, and creating auralisations and listening tests as the basis for subsequent psychoacoustic and perceptual evaluation of the acoustic quality of the investigated hall. Psychoacoustic analysis can provide an insight into additional parameters such as sharpness, fluctuation strength, tonality, or even to the perceptual values of envelopment or loudness decay, which in this hall, not being adjacent to outdoor spaces, certainly contribute to perceived reverberance. Such an approach may lead to reconstruction and interpretation of acoustic preferences of specific cultures and/or historical periods, as well as understanding of ancient spaces as multisensory, embodied environments rather than purely geometric analysis.
Author Contributions
Conceptualization, M.N.M., M.H., and Z.V.; methodology, M.N.M. and M.H.; software, M.H.; validation, M.N.M., M.H., and Z.V.; formal analysis, M.H., M.N.M., and Z.V.; investigation, M.N.M. and M.H.; resources, M.N.M. and M.H.; data curation, M.N.M., M.H., and Z.V.; writing—original draft preparation, M.N.M., M.H., and Z.V.; writing—review and editing, M.N.M., M.H., and Z.V.; visualisation, M.N.M. and M.H.; supervision, M.H. and M.N.M.; project administration, M.N.M.; funding acquisition, M.N.M. and Z.V. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
The raw data supporting the conclusions of this article will be made available by the authors on request.
Acknowledgments
The authors would like to thank the late Katja Marasović, for useful discussions, information about the cellars of Diocletian’s palace, coordination of site access, and connection to Miljenko Žabčić, who unconditionally provided the drawings of the Cellars.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Chourmouziadou, K.; Kang, J. Acoustic evolution of ancient Greek and Roman theatres. Appl. Acoust. 2008, 69, 514–529. [Google Scholar] [CrossRef] [Scilit]
- Till, R. Sound Archaeology: A Study of the Acoustics of Three World Heritage Sites, Spanish Prehistoric Painted Caves, Stonehenge, and Paphos Theatre. Acoustics 2019, 1, 661–692. [Google Scholar] [CrossRef] [Scilit]
- Rindel, J.H. Roman theatres and revival of their acoustics in the ERATO project. Acta Acust. 2013, 99, 21–29. [Google Scholar] [CrossRef] [Scilit]
- Đorđević, Z.; Penezić, K.; Dimitrijević, S. Acoustic Vessels as an Expression of Medieval Music Tradition in Serbian Sacred Architecture. Musicology 2017, 22, 105–132. [Google Scholar] [CrossRef] [Scilit]
- Vitruvius, M. Vitruvius: The Ten Books on Architecture; Morgan, M.H., Translator; 1st century BC; Harvard University Press: Cambridge, MA, USA, 1914. [Google Scholar]
- Suárez, R.; Sendra, J.J.; Alonso, A. Acoustics, liturgy and architecture in the early Christian church. From the domus ecclesiae to the basilica. Acta Acust. 2013, 99, 292–301. [Google Scholar] [CrossRef] [Scilit]
- Suárez, R.; Alonso, A.; Sendra, J.J. Archaeoacoustics of intangible cultural heritage: The sound of the Maior Ecclesia of Cluny. J. Cult. Herit. 2016, 19, 567–572. [Google Scholar] [CrossRef] [Scilit]
- van Tonder, C.; Yan, R.; Tronchin, L. Pompeii Performance Soundscapes in the Amphitheater, the Grand Theater, and the Odeon. Heritage 2025, 8, 196. [Google Scholar] [CrossRef] [Scilit]
- Berardi, U.; Iannace, G. The acoustic of Roman theatres in Southern Italy and some reflections for their modern uses. Appl. Acoust. 2020, 170, 107530. [Google Scholar] [CrossRef] [Scilit]
- Girón, S.; Galindo, M.; Romero-Odero, J.A.; Alayón, J.; Nieves, F.J. Acoustic ambience of two roman theatres in the Cartaginensis province of Hispania. Build. Environ. 2021, 193, 107653. [Google Scholar] [CrossRef] [Scilit]
- Muth, S. Historische Dimensionen des Gebauten Raumes—Das Forum Romanum als Fallbeispiel; De Gruyter: Berlin, Germany, 2014. [Google Scholar] [CrossRef] [Scilit]
- Kopij, K.; Pilch, A. The Acoustics of Contiones, or How Many Romans Could Have Heard Speakers. Open Archaeol. 2019, 5, 340–349. [Google Scholar] [CrossRef] [Scilit]
- Kopij, K.; Pilch, A.; Drab, M.; Popławski, S. One, Two, Three! Can Everybody Hear Me? Acoustics of Roman Contiones. Case Studies of the Capitoline Hill and the Temple of Bellona in Rome. Open Archaeol. 2023, 9, 20220330. [Google Scholar] [CrossRef] [Scilit]
- Bevilacqua, A.; Fuchs, W. Digital Soundscape of the Roman Theatre of Gubbio: Acoustic Response from Its Original Shape. Appl. Sci. 2023, 13, 12097. [Google Scholar] [CrossRef] [Scilit]
- Díaz-Andreu, M.; da Rosa, N.S. Exploring Ancient Sounds and Places: Theoretical and Methodological Approaches to Archaeoacoustics; Oxbow Books: Oxford, UK, 2024. [Google Scholar] [CrossRef] [Scilit]
- ISO 3382-1; Acoustics-Measurement of Room Acoustic Parameters—Part 1: Performance Spaces. ISO: Geneva, Switzerland, 2009.
- Amadasi, G.; Bevilacqua, A.; Iannace, G.; Trematerra, A. The Acoustic Characteristics of Hellenistic Morgantina Theatre in Modern Use. Acoustics 2023, 5, 870–881. [Google Scholar] [CrossRef] [Scilit]
- White, L.M. Building God’s House in the Roman World: Architectural Adaptation Among Pagans, Jews, and Christians; The Johns Hopkins University Press: Baltimore, MD, USA, 1990. [Google Scholar]
- Duran, S.; Chambers, M.; Kanellopoulos, I. An Archaeoacoustics Analysis of Cistercian Architecture: The Case of the Beaulieu Abbey. Acoustics 2021, 3, 252–269. [Google Scholar] [CrossRef] [Scilit]
- Tavares, M.A.S.M.P.; Bray, W.R.; Simmons, R. Exploring triangulation between psychoacoustic (soundscape) indicators and binaural metaphysical perceptions for music in Mission Concepcion Church, Texas. Int. Symp. Music Room Acoust. 2025, 58, 032001. [Google Scholar] [CrossRef] [Scilit]
- Martellotta, F. Understanding the acoustics of Papal Basilicas in Rome by means of a coupled-volumes approach. J. Sound Vib. 2016, 382, 413–427. [Google Scholar] [CrossRef] [Scilit]
- Cirillo, E.; Martellotta, F. Sound propagation and energy relations in churches. J. Acoust. Soc. Am. 2005, 118, 232–248. [Google Scholar] [CrossRef] [Scilit]
- Magrini, U.; Magrini, A. Measurements of acoustical properties in Cistercian abbeys. Build. Acoust. 2005, 12, 255–264. [Google Scholar] [CrossRef] [Scilit]
- Poldrugovac, T.; Horvat, M.; Vukadin, D.R. Hearing a Sacred Space: An Archaeoacoustic Analysis of the Church of St. Francis in Pula, Croatia. Acoustics 2026, 8, 16. [Google Scholar] [CrossRef] [Scilit]
- Bellia, A. Towards a digital approach to the listening to ancient places. Heritage 2021, 4, 2470–2480. [Google Scholar] [CrossRef] [Scilit]
- Gaborek, R.M. Seeing the Pompeian Domus: Soundscape Analysis at the House of the Vestals. Master’s Thesis, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA, 2022. Available online: https://cdr.lib.unc.edu/concern/dissertations/f7623p44j (accessed on 28 March 2026).
- IEC 60268-16; Sound System Equipment—Part 16: Objective Rating of Speech Intelligibility by Speech Transmission Index. IEC: Geneva, Switzerland, 2020.
- Long, M. Architectural Acoustics; Elsevier Academic Press: San Diego, CA, USA, 2006. [Google Scholar]
- Marasović, J.; Marasović, T.; McNailly, S.; Wilkes, J. Istraživanje Jugoistočnog Dijela Dioklecijanove Palače, 1 dio (1968–1971) Research of the Southeastern Part of Diocletian’s Palace, Part 1 (1968–1971), Urbs. 1972. Available online: https://www.academia.edu/70946088/01_DIOCLETIANS_PALACE_JOINT_EXCAVATIONS_I (accessed on 22 November 2025).
- Marasović, J.; Marasović, T. Diocletian Palace, Zora—Zagreb, Zagreb, 1970. Available online: https://www.academia.edu/24429864/Diocletians_palace (accessed on 22 November 2025).
- Ciaburro, G.; Berardi, U.; Iannace, G.; Trematerra, A.; Puyana-Romero, V. The acoustics of ancient catacombs in Southern Italy. Build. Acoust. 2021, 28, 411–422. [Google Scholar] [CrossRef] [Scilit]
- Iannace, G.; Trematerra, A.; Qandil, A. The acoustics of the catacombs. Arch. Acoust. 2014, 39, 583–590. [Google Scholar] [CrossRef] [Scilit]
- Zevi, B. Architecture as space. In How to Look at Architecture; Da Capo Press: New York, NY, USA, 1993. [Google Scholar]
- Marasović, J.; Marasović, T.; Gabričević, B. Research and Reconstruction of Diocletians Palace Peristyle in Split 1956–1961, Književni Krug Split, Split. 2014. Available online: https://www.academia.edu/24394002/RESEARCH_AND_RECONSTRUCTION_OF_DIOCLETIANS_PALACE_PERISTYLE_1956_1961 (accessed on 22 November 2025).
- Historical Complex of Split with the Palace of Diocletian, UNESCO World Herit. Cent. 1992–2026. Available online: https://whc.unesco.org/en/list/97/ (accessed on 26 March 2026).
- Marasović, J.; Marasović, D. Rehabilitation of the Historic Core of Split 2, City of Split: Agency for the Historic Core, Split, 1998. Available online: https://www.academia.edu/24394221/Rehabilitation_of_the_historic_core_of_Split_2 (accessed on 22 November 2025).
- Marasović, K.; Perojević, S.; Karač, Z.; Trifunović, B. Arhitekt Jerko Marasović: Renesansni Čovjek Dvadesetog Stoljeća, Architect Jerko Marasović—Renaissance Man of the Twentieth Century, Društvo Arhitekata Splita, Arhitektonski fakultet Sveučilišta u Zagrebu, 2024. Available online: https://www.arhitekti-hka.hr/hr/novosti/arhitekt-jerko-marasovi%C4%87---renesansni-%C4%8Dovjek-dvadesetog-stolje%C4%87a---predstavljanje,4807.html (accessed on 24 November 2025).
- McNailly, S.; Marasović, J.; Marasović, T. Istraživanje Jugoistočnog Dijela Dioklecijanove Palače, 2 dio (1972–1977) Research of the southeastern part of Diocletian’s Palace, Part 2 (1972–1977), Urbanistički zavod Dalamacije—Split, University of Minnesota, 1977. Available online: https://www.academia.edu/70948295/02_DIOCLETIANS_PALACE_JOINT_EXCAVATIONS_II (accessed on 22 November 2025).
- Žabčić, M. (Department of Mathematics, Boyd Research and Education Center, University of Georgia, Athens, GA, USA). Personal communication, 2023.
- ARTA Software. Available online: https://artalabs.hr/ (accessed on 24 November 2025).
- ODEON A/S. ODEON Room Acoustics Software, Version 18.0; ODEON A/S: Lyngby, Denmark, 2023. [Google Scholar]
- ODEON A/S. ODEON User Manual, Version 18; ODEON A/S: Lyngby, Denmark, 2023. Available online: https://odeon.dk/ (accessed on 24 February 2026).
- Michael, V. Auralization: Fundamentals of Acoustics, Modelling, Simulation, Algorithms and Acoustic Virtual Reality; Springer: Aachen, Germany, 2007. [Google Scholar]
- HRN DIN 18041:2024; Akustička Kvaliteta Prostorija—Zahtjevi, Preporuke i Upute za Planiranje (DIN 18041:2016 DIN 18041:2016 Hörsamkeit in Räumen—Anforderungen, Empfehlungen und Hinweise für die Planung). HZN4You Repository, HZN (Croatian Standards Institute): Zagreb, Croatia, 2016.
- Postma, B.N.J.; Katz, B.F.G. Creation and calibration method of acoustical models for historic virtual reality auralizations. Virtual Real. 2015, 19, 161–180. [Google Scholar] [CrossRef] [Scilit]
- Savioja, L.; Svensson, U.P. Overview of geometrical room acoustic modeling techniques. J. Acoust. Soc. Am. 2015, 138, 708–730. [Google Scholar] [CrossRef] [Scilit]
- Bevilacqua, A.; Farina, A.; Iannace, G.; Ferrari, J. Acoustic Investigations of Two Barrel-Vaulted Halls: Sisto V in Naples and Aula Magna at the University of Parma. Appl. Sci. 2025, 15, 5127. [Google Scholar] [CrossRef] [Scilit]
- Bradley, J.S.; Reich, R.; Norcross, S.G. A just noticeable difference in C50 for speech. Appl. Acoust. 1999, 58, 99–108. [Google Scholar] [CrossRef] [Scilit]
- NS 8178:2014; Acoustic Criteria for Rooms and Spaces for Music Rehearsal and Performance. Standard Norge: Oslo, Norway, 2014.
- Rindel, J.H. New Norwegian standard on the acoustics of rooms for music rehearsal and performance. In Proceedings of the Forum Acusticum Krakow, Krakow, Poland, 7–12 September 2014. [Google Scholar]
- Olsen, J.G. ISO 23591 Acoustic quality criteria for music rehearsal rooms and spaces. In Proceedings of the 24th International Congress on Acoustics, Gyeongju, Republic of Korea, 24–28 October 2022. [Google Scholar]
- ISO 23591:2021; Acoustic Quality Criteria for Music Rehearsal Rooms and Spaces. ISO: Geneva, Switzerland, 2021.
- Klapa Multipart Singing of Dalmatia, Southern Croatia, Represent. List Intang. Cult. Herit. Humanit. Available online: https://ich.unesco.org/en/RL/klapa-multipart-singing-of-dalmatia-southern-croatia-00746 (accessed on 26 March 2026).
- Iannace, G.; Berardi, U. Acoustic virtual reconstruction of the Roman theater of Posillipo, Naples. Proc. Meet. Acoust. 2017, 30, 015011. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.















