4.1. Overall Assessment of Suitability of the Hall for Speech
In view of reverberation time as the basic criterion for evaluation, the average mid-frequency reverberation time of 0.97 s obtained for the occupied hall almost perfectly matches the optimal reverberation time of 0.96 s for a type-A2 space of this size (volume) used for one-way speech communication, calculated according to [
44]. As shown above, the reverberance of the hall would be excessively high for low or no occupancy conditions.
The frequency-dependent values of reverberation time reveal excessive low-frequency reverberance compared to the rest of the frequency range in the unoccupied state of the hall, and the issue is even more pronounced in the occupied hall, due to the audience contributing to the total absorption in the hall predominantly at middle and high frequencies (in octave bands at 500 Hz and above). The frequency profile of reverberation time in the occupied hall exceeds the upper tolerance limit in the low-frequency region (octave bands at 125 Hz and 250 Hz), while in the unoccupied hall this discrepancy is extended to the mid-frequency region (octave bands at 500 Hz and 1000 Hz) as well. By contemporary criteria defined in [
44], the described imbalance would normally call for acoustic treatment measures focused on reducing low-frequency reverberance, thus obtaining a more balanced frequency dependence of reverberation time. However, the investigated hall is a part of Diocletian’s palace complex as a World Heritage site protected by UNESCO [
35] and as such is regarded as heritage of the highest rank. In Croatia, such sites are highly protected by relevant legislation and are exempt from requirements imposed on buildings, including the ones regarding acoustics. In this light, no permanent acoustic treatment in the investigated hall would be permitted. However, portable acoustic elements would most likely be allowed as a viable solution.
Although not as pronounced, an imbalance between mid-frequency reverberance and high-frequency reverberance can be observed in the frequency profile of reverberation time. In the occupied hall the imbalance is mitigated by the audience that introduces the most additional absorption in the mid-frequency region.
The overall judgement is that the obtained frequency profile of reverberation time is not suitable for speech by contemporary criteria due to excessive low-frequency reverberation, although the mid-frequency single-number value obtained for the fully occupied hall indicates that the amount of reverberation is optimal for the hall of this size and purpose. While spaces intended for music performances are usually designed with moderately emphasised low-frequency reverberance associated with the perception of warmth, the same is tolerated in spaces used for speech, although balanced reverberance across the low- and mid-frequency range is preferred.
In terms of speech clarity, the introduction of the audience into the hall yields significant average improvement of speech clarity of approximately 3 dB. For the unoccupied hall, the averaged single-number value of C50,mid is −2.5 dB, and negative values of C50 are observed in all individual octave bands except the ones at 4000 Hz and 8000 Hz, indicating the dominance of reverberant sound over the direct sound that has a detrimental effect on speech intelligibility. In a fully occupied hall, the averaged single-number value of C50,mid is 1.1 dB, and negative values of C50 are now observed only at low frequencies, i.e., in octave bands at 125 Hz and 250 Hz. The introduction of the audience increases the total absorption in the hall, leading to the shortening of the reverberation tail and the decrease of the energy contained in the late part of the impulse response. As the highest quantity of additional absorption is introduced in the mid-frequency range, the improvement of C50 is the greatest in this frequency range and reaches 3.5 dB. The increase at low frequencies does not exceed 2 dB, and the improvement at high frequencies is reflected in the 2.5-dB increase on average.
The spatial distribution of
C50,mid shows that the frontal part of the audience area will benefit from higher speech clarity than the rear part due to its proximity to the sound source, regardless of the level of occupancy. However, in the frontal audience area being close to the source, the change of
C50 exceeds one JND, depending on the distance from the source, whereas in the rear area its values are essentially constant, i.e., the variation is maintained within one JND, as confirmed by the summary statistics shown in
Table 5. The summary statistics also show that the difference in mid-frequency speech clarity in the unoccupied and the occupied hall is statistically significant by any criteria, as there is no overlap between the two corresponding distributions of
C50,mid. The improvement of speech clarity is in the range of 4 JND, as defined in [
16], and is, therefore, clearly perceivable.
Regarding the physical characteristics of the hall itself and their contribution to speech clarity, the hall is, in fact, rather small, quite narrow (only 7 metres wide), barrel-vaulted, and constructed of (at least nominally) acoustically reflective materials. All these characteristics should normally lead to strong early reflections as an aid to direct sound that contributes to enhanced speech clarity [
47]. However, as the hall is elongated and has parallel lateral walls, the sound coming from the source positioned in front of the apse will be strongly reflected only from the front part of these walls close to the source itself, thus benefiting only the front audience area. For the rear audience area, the angle of incidence of sound is large (approaching grazing incidence), thus causing the early reflections to miss most of the rear audience area as defined in the model. In addition, the presence of openings and passages to adjacent spaces in the lateral walls causes a part of these early reflections to be lost as they enter these openings and propagate to adjacent halls. As for the ceiling, its barrel-vaulted shape would nominally be ideal for concentrating the early reflections in the audience area [
47]. However, as shown in
Figure 2, the ceiling follows Vitruvius’s
stucco treatment guideline [
5] but also has suffered considerable damage over the centuries, especially its top part covered with tufa, thus having considerable absorption and diffusion. As such, it cannot provide strong early reflections in the audience. The described physical properties of the hall ultimately lead to limited early energy in the audience, and, consequently, lower speech clarity.
Despite the barely acceptable values of speech clarity for a space that is to be used for speech-based events, a basic analysis of speech intelligibility in the unoccupied hall has shown that fair speech intelligibility (
STI > 0.45) can be achieved in the entire audience area for native listeners with normal hearing even in the unoccupied hall, while a small part of the audience area closest to the speaker will experience good speech intelligibility (
STI > 0.60). In a fully occupied hall, good speech intelligibility can be achieved in the entire audience area. A more detailed analysis is based on the nominal qualification bands defined for the
STI in IEC 60268-16:2020 [
27] that transcend the five basic intelligibility categories. According to this qualification, the range of
STI from 0.64 to 0.72 achieved in the occupied hall corresponds to B (0.68 ≤
STI < 0.72) and C (0.64 ≤
STI < 0.68) qualification bands. Both bands designate high speech intelligibility appropriate for theatres, speech auditoria, parliaments, courts, and other venues where complex messages that contain unfamiliar words need to be verbally conveyed to the audience. The range of
STI from 0.55 to 0.63 achieved in the unoccupied hall falls mostly within the E (0.56 ≤
STI < 0.60) and F (0.52 ≤
STI < 0.56) qualification bands and only marginally in the D (0.60 ≤
STI < 0.64) band. The D band designates good speech intelligibility appropriate for conveying complex messages using familiar words, e.g., in lecture halls, classrooms, or concert halls. The E and F bands designate speech intelligibility appropriate for conveying complex messages within a familiar context, with the E band being appropriate, e.g., for concert halls and modern churches, and the F band being sufficient for shopping malls, VA systems, and cathedrals [
27].
The summary statistics of the
STI show that the introduction of the audience into the hall leads to a statistically significant increase in speech intelligibility, as the distributions of the
STI for the two examined occupancy conditions do not overlap. The improvement of speech intelligibility is considerable and is reflected in the increase of the
STI by 0.09 on average, i.e., in the range of 3 JND, as determined in [
48]. As analysed above, the introduction of the audience leads to improvement of speech intelligibility by two full
STI qualification bands.
The obtained values of the
STI are additionally analysed in terms of communication in a non-native language. As defined in IEC 60268-16:2020 [
27], the thresholds of the
STI for five basic levels of speech intelligibility are shifted upwards for non-native listeners, depending on their language proficiency level. In this context, experienced non-native listeners would still be able to perceive fair speech intelligibility anywhere in the audience area, regardless of occupancy, while good speech intelligibility would be achieved only in the few front rows closest to the speaker in the fully occupied hall. Non-native listeners with intermediate proficiency would experience fair speech intelligibility anywhere in the audience area in the fully occupied hall and only in 10% of the audience area closest to the speaker in the unoccupied hall. Good speech intelligibility would be impossible to achieve in this case.
In terms of audibility of speech, the plots displayed in
Figure 10 and the summary statistics in
Table 7 show that the overall A-weighted level of speech achieved in the audience area ranges from 57.9 to 62.4 dBA, indicating a high uniformity of level and, consequently, perceived loudness of the delivered spoken content. This observation is in line with the acoustic conditions in the hall dominated by reverberant sound due to the small amount of total sound absorption present in the unoccupied hall. The summary statistic for the unoccupied hall indicates that 50% of the audience area, i.e., its rear section, experiences a practically constant sound pressure level that stems from dominant reverberant sound, as the median value of
SPL(
A) is only 1.4 dB higher than the one calculated as the minimum. The other half of the audience area, i.e., the front section that is closer to the source, exhibits an increase in sound pressure level of 3 dB from the seats farthest from the source to the ones closest to it due to the increasing influence of direct sound from the sound source. Thus, in the unoccupied hall (or with low occupancy), a listener would be able to clearly hear the speech content everywhere in the audience, provided that the background noise is kept sufficiently low. However, for improved intelligibility, it would be preferable for them to sit in the front section of the audience.
Due to considerable additional absorption added to the hall by the presence of the audience, in a fully occupied hall the overall sound pressure level is reduced as a direct consequence of the reduction of the level of the reverberant component of sound. As shown in
Figure 10 and in
Table 7, the greatest reduction of approximately 4 dB is observed in the rear section of the audience located far away from the source, in which the overall level is dominated by reverberant sound. In the parts of the audience closest to the source, the observed reduction reaches only 2 dB due to the dominance of direct sound. The overall level now ranges from 53.7 to 60.2 dBA, indicating reduced uniformity of level compared to the unoccupied state. As before, the sound pressure level in the rear section of the audience, taking approximately 50% of the total audience area, is maintained within a small margin of 1.8 dB, whereas in the front section of the audience it increases by almost 5 dB as the distance between the listener and the source decreases.
In terms of both audibility of speech and uniformity of speech level, the absence of an audience (or in low occupancy conditions) inherently leads to better audibility and uniformity. On the other hand, the fully occupied hall provides more favourable acoustic conditions in terms of speech clarity and intelligibility. However, the reduction of the overall sound pressure level by 2 to 4 decibels is noticeable and easily perceivable as a reduction of the loudness of the delivered speech. Given that the audience itself can be an internal source of background noise, any reduction in speech level is detrimental to its audibility and impacts the ability of the audience to follow and comprehend the spoken content. As shown above, the rear section of the audience would suffer the greatest reduction of speech level in the occupied state of the hall, thus being more prone to the described difficulties.
The historical, architectural, and cultural significance of the hall and the entire complex of Diocletian’s palace provide a high potential for its revitalisation and possible use for a variety of events. In light of the four investigated contemporary criteria and the overall suitability of the hall for speech-based events, the overall judgement is that the hall in its present condition can be used for speech-based events that (1) predominantly employ one-way communication between the speaker and the audience, (2) assume high occupancy of the audience area, and (3) imply a low level of background noise in the hall, as well as a trained speaker capable of producing a sufficiently high speech level. As a few examples, the events befitting this hall could include poetry readings or literary recitals in general, as well as various sorts of presentations, meetings, openings of exhibitions, etc. However, for two-way verbal communication involving multiple speakers, e.g., mingling after the official part of an event has ended, the hall would need to undergo additional acoustic treatment to further reduce its broadband reverberance, on top of dealing with excessive low-frequency reverberance, as suggested above. Current conditions in the hall in terms of speech intelligibility would be suitable for events being held in a native language, whereas the intelligibility of speech delivered in a non-native language would be somewhat compromised. Given that the Diocletian’s palace complex and the city of Split in general receive a high number of foreign visitors each year, further improvement in speech intelligibility in the hall would facilitate the extension of the events happening in the hall to the foreign tourist population as well. As permanent acoustic treatment would be regarded as invasive in this kind of space, intelligibility could be improved by adding a sound reinforcement system composed of loudspeakers with higher directivity than the one exhibited by a live speaker. Such an addition would simultaneously (1) improve the direct-to-reverberant ratio, thus improving clarity and intelligibility, (2) improve speech-to-noise ratio in elevated background noise conditions, again leading to improved intelligibility, and (3) provide aid for untrained speakers with limited vocal ability (in terms of volume).
4.2. Overall Assessment of Suitability of the Hall for Music
In terms of reverberation, the results indicate that the reduction of the audience area from 135 seats for speech to 81 seats for music performances was insufficient to achieve optimal reverberance for music. In the unoccupied hall, the mid-frequency reverberation time exceeds the optimal value but is still within the defined tolerance. Thus, further reduction of the audience area would be needed to achieve a match of the calculated value and the one found to be optimal according to [
44]. However, such a reduction would have to be significant, reducing the audience to 20 to 30 seats at most, which raises the question of the feasibility of music-based events, thus making such a step unlikely.
Viewed as a frequency-dependent parameter, the reverberation time again exhibits excessive low-frequency reverberance, leading to a perception of warmth provided by the hall in terms of acoustics, and the presence of the audience further emphasises the issue. Similar to the setup of the hall made for speech, the hall in the setup for music again displays lower high-frequency reverberance compared to the one observed in the mid-frequency region, thus leading to a reduced perception of brilliance and liveness and further adding to the perception of warmth.
Figure 11 shows that the frequency profile of reverberation time suffers from a lack of uniformity. For the occupied hall, it is kept within the tolerance limits only in the three lowest octave bands (at 125 Hz, 250 Hz, and 500 Hz), whereas it falls below the tolerance zone in the rest of the frequency range. In the unoccupied hall the inconsistency of reverberation time is taken to the extreme, as it meets the tolerances only in the octave band at 1000 Hz. The values in the lower bands overshoot the tolerance range, while the ones in the higher bands fall below it.
Realistically, the perception of the acoustic situation in the hall would depend on the performer and the type of instrument they play. In terms of objective criteria, the mid-frequency reverberation time in the occupied hall is found near the lower limit of the range of typical values for performance spaces defined in [
16] and is too low for music performances, according to [
44]. In addition, the obtained frequency profile of reverberation time in the occupied hall is not suitable for music performances, as the hall is not reverberant enough in the mid- and high-frequency ranges and simultaneously displays excessive low-frequency reverberation. The solution to this issue would require (1) further reduction of the audience area to achieve sufficient mid-frequency and high-frequency reverberation, and (2) introduction of low-frequency absorption into the hall to reduce low-frequency reverberation and balance it out with reverberation in the rest of the frequency range. At present, neither step is viable, as (1) it would not be feasible to hold any kind of music performance for a very small audience, and (2) no permanent acoustic interventions would likely be allowed in the hall, as explained above, but only the ones of temporary nature.
In terms of sound strength calculated for the hall, the frequency dependence of sound strength reflects the frequency profile of reverberation time. The highest values of strength are found at low frequencies, both in the unoccupied and in the occupied hall, and the sound strength monotonically decreases with the increase of frequency. The introduction of the audience into the hall results in an average decrease in sound strength of 2 dB in the mid-frequency range, i.e., in the octave bands at 500 Hz, 1000 Hz, and 2000 Hz. According to [
16], this difference is equal to 2 JND and is, therefore, perceivable. The reduction of sound strength in the rest of the frequency range is kept around 1 dB and is, therefore, just noticeable.
The spatial distributions of single-number sound strength Gmid show that the introduction of the audience into the hall results in a statistically significant decrease in sound strength, as the difference of the median values for the two cases of occupancy exceeds 3 dB, i.e., 3 JND. In terms of variation with the distance from the source, higher strength was observed in the frontal part of the audience area that is closer to the source, regardless of occupancy. In the rear half of the audience area, the values of Gmid are contained within 1 JND, suggesting that the entire rear half of the audience essentially experiences the same perception of loudness. A clear perceptual difference in terms of strength and loudness can be expected only for the 10% of the audience seated in the frontmost rows, compared to the rear half of the audience.
The suitability of rooms for different music performances and corresponding ensembles in terms of sheer loudness was assessed by considering the obtained values of strength
Gmid for the entire audience area, the size of the room in terms of its net volume
V, and reverberation in terms of reverberation time
T. In the design stage of both performance [
49,
50] and rehearsal rooms [
51,
52], these quantities and the relationships between them are adjusted so that the room fits the performers (a soloist, a small ensemble, or a large ensemble) in terms of the level of the played sound. In the frame of this study, all these values are known for the investigated hall and are not subject to change, as no permanent acoustic intervention (or intervention of any kind) is allowed in the hall.
Given the volume of the hall of 950 m
3, its floor area of 156 m
2, and the size of the audience area of 81 seats, it is reasonable to assume that the hall would be suitable for music performed by a soloist or a small ensemble at most. According to the procedure described in the ISO 23591 standard [
52], to achieve the target level at
forte between 85 dB and 90 dB, and having in mind the average mid-frequency sound strength
Gmid of 13.5 dB in the fully occupied hall, the sound power level of a soloist or an ensemble should be in the range from 102.5 dB
re1pW to 107.5 dB
re1pW, which translates to sound power in the range from 18 mW to 56 mW.
An example of a small ensemble that could be performing in the hall is the Dalmatian
klapa, traditionally consisting of anywhere from 4 to 10 male singers, singing a capella, inscribed in 2012 on the Representative List of the Intangible Cultural Heritage of Humanity [
53]. In recent decades, all-female and mixed klapas are present as well. Considering the average size of a klapa of 7 singers and taking the sound power data from [
52], the total sound power of an average-sized klapa can be set to 25 mW, thus making the ensemble capable of reaching the forte level of 86.5 dB in the audience area.
A classic string quartet consisting of two violins, a viola, and a cello reaches the total sound power of 3.1 mW, thereby producing the level of only 77.4 dB in the audience area, which proves to be insufficient for the audience to perceive the loudness appropriate for playing at true forte but is more appropriate for an intimate performance.
A solo singer accompanied by a guitar is another traditional music form in Dalmatia. Together, these two sound sources yield the total sound power of 4.4 mW, which corresponds to the sound pressure level of 79 dB. However, the audience would likely expect a lower overall level from such a performer and a performance of such an intimate nature.
A solo trumpet has the sound power of 12.5 mW and, by the same calculations, would produce the sound pressure level in the audience of 83.5 dB, just shy of the lower limit set for optimal levels at forte. However, the calculations presented in [
52] do not consider source directivity but are limited to the level of reverberant sound achievable in a room characterised by strength
G and excited by an omnidirectional sound source with a known sound power. For highly directional sources such as the trumpet, the audience in this hall is expected to be in the direct sound field of the instrument, thus being exposed to direct sound, the level of which greatly exceeds the one calculated for reverberant sound.
Finally, even a small ensemble of three or four loud brass instruments would already exceed the sound power limit calculated above for this particular hall, proving to be excessively loud while playing at forte in this hall.
Similar to speech clarity calculated for the hall set up for speech-based events, a significant average improvement of mid-frequency music clarity of approximately 2.5 dB can be expected with the audience present in the hall. The frequency-dependent values of the spatially averaged music clarity shown in
Figure 14 for the unoccupied hall reveal that negative values of
C80 are limited to octave bands at 500 Hz and below, indicating the dominance of reverberant sound over the direct sound. With the hall fully occupied, a negative value of
C80 is now observed only in the octave band at 125 Hz. Due to the sound absorption properties of the audience, the increase of
C80 is again the greatest in the mid-frequency range, as stated above. The increase at low (125 Hz and 250 Hz) and high (4000 Hz and 8000 Hz) frequencies is kept below 2 dB. In terms of the typical range of music clarity defined in [
16], i.e., from −5 dB to +5 dB for performance spaces, the median music clarity in the unoccupied hall is in the midpoint of the defined range at 0 dB, implying a balance between direct and reverberant sound. In the occupied hall, the same value reaches 2.5 dB, thus marking a shift towards clearer, but simultaneously drier sound.
Figure 15 shows the spatial distribution of the single-number mid-frequency music clarity
C80,mid. Compared with clarity calculated for speech, some commonalities can be found. As found for speech, higher music clarity was found in the frontal part of the audience area, as it is closer to the source, both in the unoccupied and in the occupied hall. However, the reconfiguration of the audience area to a single section that is moved away from the sound source to the central part of the hall is reflected in the obtained values of
C80,mid. While the range of
C50,mid found in the audience area for speech was 4 dB in the unoccupied hall and 3.2 dB in the occupied hall, the
C80,mid calculated for the entire one-section audience is kept within a 1.7-dB range in the unoccupied hall and within a 1.6-dB range in the occupied hall, just exceeding 1 JND in both cases. This suggests that the perceived clarity of music should be quite similar for all persons seated in the audience area. As with speech clarity, the difference in mid-frequency music clarity for the two extreme cases of occupancy is statistically significant by any criteria, as the summary statistics show no overlap between the two corresponding distributions of
C80,mid. The improvement of music clarity is in the order of 2.5 JND, as defined in [
16], thus making it clearly perceivable.
In terms of the early lateral energy fraction, its single-number value proves to be quite high. The median value of 0.290 was obtained for the unoccupied hall and 0.278 in the occupied hall, while individual values exceed 0.250 in the entire audience area. The typical range of values for performance spaces defined in [
16] is from 0.100 to 0.350. In this context, the hall is expected to provide strong early lateral energy to the audience, thus contributing to the perceptive broadening of the sound source. High values of this parameter are attributed to the overall shape of the hall, as the ratio of length to width is approximately 3 to 1, and the width of the hall in absolute terms is only 7 metres. In addition, all the walls of the hall, including the lateral ones, are highly reflective as well, thus producing strong lateral reflections that arrive to the audience shortly after direct sound. The summary statistics reveal that the
JLF,average is maintained within one JND of 0.05, as defined in [
16], in the entire audience area, in both the unoccupied and the occupied hall. Moreover, the influence of occupancy on
JLF,average is marginal, as the reduction caused by introducing the audience is far below the defined JND.
The spatial distributions of
JLF,average shown in
Figure 16 reveal interesting, localised effects of the layout of lateral walls on the
JLF,average in the audience area, i.e., the front right part of the audience area displays higher values of
JLF,average than observed in the front left part. The difference is minute and has been attributed to the position of the sound source and the audience area in the hall, as well as to the layout of the passages in both lateral walls towards adjacent spaces. These openings are essentially “black holes” for early reflections, as any sound that enters a passage will suffer multiple reflections from passage walls and ultimately travel into the adjacent space to which the passage leads. At first glance, the right wall has two passages in the critical area that provides lateral reflections for the audience, whereas the left wall has only one. Thus, the early lateral energy fraction was expected to be higher in the left part of the audience. However, additional analysis was performed in terms of calculating the spatial distributions of
JLF,average for five alternative positions of the sound source chosen along the width of the hall. In addition, the paths of the first reflections from the lateral walls were traced from the source position to the audience area in each of the cases, including the original one. For the original source position, it was ultimately concluded that there are two clear paths for the first reflections from the lateral walls to the front right part of the audience, as the incident sound meets a flat surface on both lateral walls to be reflected from them into the audience. On the other hand, only one early reflection path has been identified for the front left part of the audience area, i.e., the first reflection arrives only from the left lateral wall, as the right lateral wall contains a passage exactly in the area sound would reflect from, thus eliminating this early reflection.
As discussed above, the value of the hall in terms of architecture and heritage and the fact that the entire complex of Diocletian’s palace is a unique heritage site should be the driving force for the use of the hall as a venue for music performances as well. Given that Dalmatia has a great tradition in group a cappella singing, a logical step would be to use the hall as a venue for klapa performances. The experience of listening to this kind of performance in a historically and architecturally unique space such as this one would undoubtedly be rewarding for a variety of audiences, both domestic and foreign. In the context of the investigated contemporary criteria and the overall assessment of the suitability of the hall for music performances, the hall in its current condition, with reduced audience area, is nominally not suitable for music performances in the occupied condition, as mid-frequency reverberation in the hall is shorter than required, while low-frequency reverberation is excessive. The gain provided by the hall and its overall size suggest that the hall can and should serve as a venue for solo or small-ensemble music performances. Positive values of music clarity in the occupied hall point to a clear sound provided by the hall in terms of the possibility for the listener to distinguish individual notes even in fast passages. Due to the size and shape of the hall, the size of the audience area, and the presumed types of music performances to take place inside, the obtained values of early lateral energy fraction are rather high. Small-hall performances lean towards clarity, intimacy, and communication between the performers and the audience, rather than an unnatural and overemphasized perceived width of the sound source. Given all the factors mentioned above, in this particular hall the perception of an excessively large sound source might disagree with its actual size and with the overall acoustic information received by a listener seated in the audience.
However, in the context of klapa singing, excessive low-frequency reverberance would provide a warmer tone to the entire ensemble by enhancing the baritone and, especially, bass voices that are often in a secondary role and lower in volume compared to the rest of the ensemble. As discussed above, the size of the hall and its reverberance yield a gain that is quite suitable for a small ensemble of singers such as a klapa. The clarity of sound in the hall would enable the listeners to clearly understand the lyrics of the song, as an important aspect of Dalmatian tradition, and perhaps even to sing along if appropriate. In terms of perceptual broadening of the sound source, in the case of klapa, it would not disagree with the actual size of the ensemble, as the singers are normally placed in a single row or arc and would occupy a substantial part of the actual width of the hall (depending on the number of singers).
4.3. Comparison with the Results of Similar Archaeoacoustic Studies
As the final step of the analysis, a comparison was made of the obtained results for the investigated hall with the results of similar archaeoacoustic studies. The greatest challenge in this analysis was to find a study performed on a heritage site similar to the one examined in this study. As the investigated hall is quite unique, this challenge has not been successfully overcome. Instead, a comparison of the investigated site was made with a wide range of heritage sites, ranging from prehistoric caves to ancient Greek and Roman theatres, catacombs, and early Christian churches.
As presented in previous sections, the investigated hall was evaluated regarding its potential use as a space for both speech and music. The values of the key parameters that were obtained for the speech configuration are: T20,mid = 1.39 s (unoccupied) and 0.97 s (occupied), C50,mid = −2.5 dB (unoccupied) and 1.1 dB (occupied), STI = 0.55–0.63 (unoccupied) and 0.64 to 0.72 (occupied), and SPL(A) = 58–62 dBA (unoccupied) and 54–60 dB (occupied). The values of the key parameters that were obtained for the music configuration are: T20,mid = 1.53 s (unoccupied) and 1.16 s (occupied), C80,mid = 0.6 dB (unoccupied) and 3.0 dB (occupied), Gmid = 15.5 (unoccupied) and 13.5 (occupied), and JLF,average = 0.277–0.317 (unoccupied) and 0.262–0.306 dB (occupied).
The type of heritage sites that are extensively analysed are the open-air theatres from the Greek and Roman eras. In [
1], an extensive analysis was made in terms of the evolution of open-air theatres from the early era of Greek civilisation to Roman times, in terms of their construction and, consequently, acoustic properties. It was found that theatre construction has evolved with time in terms of increasing the level of speech in the audience area and lengthening the reverberation time. In terms of specific values, the obtainable speech level in the audience area has been increased to 53 dB at the point closest to the stage, i.e., 25 m away from the source. As the distance increases (to 65 m as the maximum), the level drops in a manner quite similar to the one observed in the free field. However, to achieve the stated levels, the sound pressure level of the source was set to “performance level”, i.e., considerably louder than speaking in a normal or even raised voice. The reverberation time was found in the range from 0.3 s for early designs to 1.5 s in Roman theatres. As for speech intelligibility, being an important feature for theatres, the speech transmission index
STI was found to be in the range from 0.42 (poor) to as high as 0.80 (excellent). Specifically for Roman theatres, the values of the
STI range from 0.45 to 0.53 in the unoccupied state and from 0.51 to 0.70 in the occupied setup, suggesting that the acoustics of Roman theatres is similar to the one aimed for in modern ones.
As the design of the investigated hall is fundamentally different from an open-air theatre, the comparison of the two designs is challenging, as their acoustic behaviour is quite different as well. As outdoor spaces, open-air theatres lack the late reflections and true reverberation tail in their impulse response, which is composed predominantly of direct sound and early reflections. Hence, they are characterised by a reasonably short reverberation time, high speech clarity and intelligibility, and non-uniform speech level distribution in the audience area. As a specific example, the Roman theatre of Gubbio investigated in [
14] displays very high speech and music clarity in the fully open state, ranging from 13–14 dB and from 14–18 dB, respectively. Both parameters far exceed the values found in the investigated hall. On the other hand, the sound strength in the theatre is quite low (around 0 dB on average), as expected for a fully open space. An architectural intervention in the form of adding the upper porch and/or roof to partially close the otherwise open theatre led to a considerable reduction of speech clarity to the range from −2 to 3 dB and music clarity from 2 to 8 dB, thus being more in line with expected values for a space of this type. Simultaneously, the values of sound strength in the audience area increased to the range from 4 dB to 10 dB, thus contributing to the audibility of the spoken content. As another example of a Roman theatre investigated in [
54], the theatre in Posillipo, Naples, displays the value of reverberation time
T30,mid of approximately 1.3 s, while its
STI values ranging from 0.58 to 0.75 are similar to those obtained in the investigated hall, as well as the values of music clarity of 1.5 dB to 5 dB. An investigation presented in [
2] focuses on the Paphos theatre. The values of reverberation time of 1.75 s and RASTI ranging from 0.50 to 0.55 are again similar to those obtained in the investigated hall.
Enclosed spaces such as early Christian basilicas, catacombs, or even prehistoric caves are considerably more similar to the investigated hall in terms of the way the overall sound field forms inside the space. Yet, in terms of sheer size and shape, a suitable counterpart to the investigated hall has not been found in literature. In [
2], a study of selected prehistoric caves has been made, and it was determined that the reverberation time inside is found in the range from 0.75 to 1.50 s, while the speech transmission index expressed as RASTI was found to be in the range from 0.62 to 0.77. As investigated in [
31], the upper level of the Catacombs of San Gennaro, Naples, proved to be quite similar to the investigated hall in terms of shape (elongation) and stone cladding but considerably larger, of a more developed floor plan and including also porous earthen surfaces filled with burial niches. Nevertheless, the values of
T30,mid of 1.14 s, and the
STI of 0.63 mark this space as acoustically quite comparable to the investigated hall, along with the caves from the previous example. Similar values of reverberation time and speech transmission index in the catacomb and the investigated hall, despite a considerable size difference, point to a larger amount of total sound absorption in the catacomb attributed to a large number of earthen surfaces and the shaping of the walls through rows of burial niches.
Early Christian churches, researched in [
6] in terms of acoustics and architectural typology, inherited the form of Roman basilicas, representing a potential reference for comparison with the investigated single-nave hall. An early domus with a small room volume of only 340 m
3 was compared with larger basilicas (up to 75,000 m
3), showing the values of
T30,mid of around 1.0 s and
STI 0.60, to be comparable to Roman and early Christian contexts, where speech intelligibility was often prioritised. Larger basilicas reflect later liturgical developments exhibited with long reverberation appropriate for ritual vocalisation of the liturgical chant [
6,
7,
21], accompanied by poor speech intelligibility. As such, they are in no way comparable with the investigated hall or with other discussed types of heritage sites.