1. Introduction
When an operator decides to deploy or expand an Over-The-Top (OTT) or Internet Protocol Television (IPTV) service, one of the earliest and most critical infrastructure decisions concerns the organisation of server-side delivery resources. Concretely, the operator must choose between at least three distinct provisioning architectures: assigning dedicated server pools to each subscription tier, aggregating all traffic classes into a single shared pool, or operating a shared pool with explicit traffic prioritisation. Each of these choices carries a different capital and operational cost profile and delivers a different level of service protection to individual user groups. Yet in the early stages of infrastructure design—before deployment trials or detailed simulation campaigns are feasible—the operator typically lacks an analytical tool to quantify these trade-offs in terms of the minimum delivery capacity required under each strategy.
The dimensioning problem is challenging because operator-managed OTT/IPTV platforms simultaneously serve multiple user populations with heterogeneous resource demands and different service-continuity requirements [
1]. Free-tier subscribers may tolerate occasional service interruptions, whereas premium subscribers are bound by contractual commitments requiring blocking probabilities several orders of magnitude lower. At the same time, resource pooling introduces statistical multiplexing gains, and traffic prioritisation can concentrate the blocking load on lower-priority classes, allowing the shared resource pool to be sized more compactly than a tier-by-tier approach would suggest. Quantifying these effects analytically is therefore of direct engineering value.
The grade-of-service (GoS) framework for server-side dimensioning is well established in teletraffic engineering [
2]. Within this framework, the key metric is the blocking probability: the probability that an arriving service request cannot be admitted because insufficient delivery capacity is available. The relationship between blocking probability, offered traffic, and server capacity in multiservice systems is characterised by the Kaufman–Roberts recursion [
3,
4], which enables computationally efficient evaluation of class-specific blocking probabilities in a full-availability group (FAG). GoS-based dimensioning complements the quality-of-experience (QoE) perspective studied in the literature [
5,
6,
7]: while QoE measures capture end-user perception of delivered video quality, GoS measures characterise the admission-level service guarantees that the delivery infrastructure can sustain.
Despite the maturity of teletraffic models in the telecommunications domain, their systematic application to OTT/IPTV server-side dimensioning has not been fully explored. The majority of existing studies address QoE-aware resource allocation [
5,
6,
7], dynamic traffic management [
8,
9], and demand prediction [
10,
11]—all of which are oriented toward online control rather than preliminary infrastructure sizing. The problem of selecting and quantifying the capacity consequences of a provisioning architecture before deployment remains insufficiently addressed by these approaches.
This paper addresses that gap directly. While the analytical tools used in this work, such as the Kaufman–Roberts recursion and the full-availability group model, are well established in teletraffic engineering, their application to OTT/IPTV server-side capacity dimensioning has not been systematically explored. In particular, no prior work has formulated this problem within a blocking-probability-based framework with differentiated service constraints across subscription tiers. The novelty of the present paper lies in the formulation of the OTT/IPTV server-side dimensioning problem in a multiservice loss-system framework, the unified analytical comparison of alternative provisioning strategies, and the translation of capacity results into concrete engineering quantities directly applicable to operator planning.
It should be emphasized that the contribution of this work is not the introduction of a new analytical model, but the formulation and solution of a practically relevant and previously unaddressed dimensioning problem using exact and well-established teletraffic tools. An analytical framework for OTT/IPTV capacity dimensioning is proposed based on full-availability group models and the Kaufman–Roberts recursion, and applied to a six-class scenario representative of a moderate-scale operator deployment. Three provisioning architectures are evaluated and compared: dedicated servers per subscription tier, a shared non-prioritised server, and a shared prioritised server. The results are interpreted not only in terms of percentage capacity savings, but also in terms of their concrete implications for infrastructure sizing—specifically, the number of additional simultaneously served streams or deployed set-top boxes that the capacity difference can sustain.
To address the identified gap in analytical OTT/IPTV capacity dimensioning, this paper makes the following contributions:
Problem formulation and analytical framework: The server-side capacity-dimensioning problem for operator-managed OTT/IPTV systems is formulated as a multi-class loss system with differentiated blocking-probability constraints reflecting subscription tiers. Based on this formulation, an operator-oriented analytical framework is developed using the full-availability group (FAG) model and the Kaufman–Roberts recursion.
Capacity-dimensioning procedure: A reproducible dimensioning method is proposed that combines recursive occupancy-distribution evaluation with an incremental capacity search procedure to determine the minimum delivery capacity satisfying class-specific grade-of-service (GoS) requirements.
Comparison of provisioning strategies: Three practically relevant resource provisioning architectures—dedicated, shared, and prioritised shared—are analysed within a unified framework. The impact of statistical multiplexing and traffic prioritisation on the required server capacity is quantified.
Per-tier blocking analysis: The framework provides a detailed per-class evaluation of blocking probabilities, enabling analysis of how congestion is redistributed across subscription tiers under different provisioning strategies, particularly in the presence of prioritisation.
Operator-oriented interpretation: The analytical results are translated into engineering quantities directly relevant to infrastructure planning, including the supported number of simultaneous streams and the corresponding set-top box (STB) population.
Model validation: The analytical model is validated using discrete-event simulation, confirming its accuracy over the considered operating range and supporting its applicability for practical dimensioning tasks.
To further clarify the key contribution of this work, the addressed research problem can be expressed through the following research questions:
RQ1: How can server-side capacity in multi-class OTT/IPTV systems be dimensioned analytically under differentiated blocking probability constraints? Answer: By formulating the system as a multi-class loss model based on a full-availability group and applying the Kaufman–Roberts recursion combined with a capacity search procedure.
RQ2: What is the quantitative impact of alternative resource provisioning strategies (dedicated, shared, prioritised) on the required delivery capacity? Answer: By evaluating the minimum capacity satisfying class-specific GoS constraints under each provisioning architecture within a unified analytical framework.
RQ3: How does traffic prioritisation influence the distribution of blocking probabilities across subscription tiers? Answer: By analysing class-specific blocking probabilities and showing how prioritisation redistributes congestion from higher- to lower-priority traffic within a shared resource pool.
The remainder of this paper is organised as follows.
Section 2 reviews the related literature.
Section 3 presents the OTT/IPTV scenario, the modelling assumptions, and the analytical framework.
Section 4 reports the validation and dimensioning results.
Section 5 discusses the engineering implications and limitations.
Section 6 summarises the conclusions.
2. Related Work
Research related to OTT and IPTV service delivery spans several complementary areas, including QoE-aware resource allocation, adaptive control, user behaviour analysis, demand prediction, and analytical teletraffic modelling. The following review identifies the gap that motivates the present work.
QoE-aware resource allocation and control in OTT systems has been studied extensively. The NOVA algorithm enables joint allocation of network resources and adaptation of video quality for multiple users, taking into account average quality, temporal variability, and fair resource sharing [
5]. Deep reinforcement learning has been applied to maximise QoE in OTT streaming over IoMT networks using the Proximal Policy Optimization method [
6]. In [
7], a QoE management system leveraging SDN and NFV was proposed for OTT services over next-generation mobile networks. The S2VC algorithm uses SDN to control resource allocation with the aim of maximising QoE metrics in HTTP Adaptive Streaming [
8]. A low-complexity heuristic has also been proposed for joint optimisation of caching, transcoding, and radio resource allocation in 5G MEC deployments [
9]. Collaborative OTT-ISP service management approaches have also been studied, addressing QoE-aware interaction between content providers and network operators [
1]. While these works provide valuable insight into QoE monitoring and optimisation, they do not address analytical server-side capacity dimensioning.
Demand modelling and traffic prediction for IPTV/OTT services have also attracted attention. Machine-learning-based predictive models were applied to forecast resource demand from Catch-Up TV usage data [
10]. User activity logs from an IPTV operator were analysed and used to develop the Simulwatch workload generator [
11]. A combinational Latent Dirichlet Allocation model was proposed to infer user preferences for real-time resource allocation in IPTV systems [
12]. These works support simulation-based dimensioning but do not provide analytical blocking-probability-based capacity estimates.
The analytical framework proposed in the present paper is rooted in classical multiservice loss models. The Erlang loss formula characterises the blocking probability in single-service systems [
2]. Its extension to heterogeneous traffic was independently established by Kaufman [
3] and Roberts [
4], who derived the product-form occupancy distribution and the Kaufman–Roberts recursion. Multiservice loss models have been extended to accommodate variable-bit-rate traffic, finite-population sources, and the Binomial–Poisson–Pascal (BPP) traffic model [
13]. Priority-based extensions of full-availability group models, in which higher-priority classes are analytically protected from lower-priority traffic, have been applied in the context of WCDMA and LTE radio interfaces [
14,
15]. Importantly, the priority mechanism used in these works is based on the general FAG framework with arbitrary integer resource demands, rather than on technology-specific properties of radio access systems. This formulation provides the theoretical basis for the priority mechanism used in the present paper. A comprehensive treatment of multirate teletraffic loss models is provided in [
16]. From the OTT/IPTV perspective, the notion of “displacement” should not be interpreted as preemptive interruption of already established streaming sessions. Instead, the mechanism operates strictly at the session-admission level. In practice, this corresponds to operator policies in which higher-priority traffic is protected through effective capacity reservation or admission thresholds, while lower-priority requests are admitted only when sufficient residual capacity is available. As a result, under high system load, lower-priority traffic experiences increased blocking, whereas premium-tier services remain largely unaffected.
Despite the maturity of analytical teletraffic models in the telecommunications domain [
2,
3,
4,
13,
16], their application to server-side capacity dimensioning of OTT/IPTV delivery systems remains limited and fragmented. At the same time, existing OTT/IPTV studies are primarily focused on QoE-aware resource allocation, adaptive control, and demand prediction [
5,
6,
7,
8,
9,
10,
11], rather than on analytical infrastructure dimensioning. In particular, three key shortcomings can be identified in the existing literature. First, there is a lack of blocking-probability-based analytical frameworks specifically tailored to server-side OTT/IPTV dimensioning, where admission control and finite delivery capacity are the primary constraints. Second, existing works do not provide a systematic and unified comparison of alternative resource provisioning strategies (dedicated, shared, and prioritised) under differentiated grade-of-service (GoS) requirements across multiple subscription tiers. Third, the available approaches do not establish a direct link between analytical capacity results and practical engineering quantities relevant to operators, such as the supported population of set-top boxes (STBs) or the number of concurrent service sessions. These limitations motivate the development of the analytical framework proposed in this paper, which addresses all three aspects within a unified and operator-oriented formulation.
3. Materials and Methods
This section presents the analytical methodology used for server-side capacity dimensioning in a multi-class OTT/IPTV delivery system. First, the considered service architecture is reduced to a tractable resource model focused on the CDN edge server cluster. Next, the system is represented as a full-availability group, and class-specific blocking probabilities are calculated using the Kaufman–Roberts recursion. The analytical framework is then validated against discrete-event simulation and applied to three provisioning strategies: dedicated, shared, and prioritised shared resource pools.
3.1. OTT/IPTV Delivery Scenario
The considered system represents an OTT/IPTV service delivery architecture in which multimedia streams are delivered to end users through a Content Delivery Network (CDN) edge infrastructure. From the perspective of resource dimensioning, the main element of interest is the server-side delivery platform responsible for supporting concurrent streaming sessions. End-user devices, such as set-top boxes (STBs) or equivalent client terminals, generate service requests that occupy delivery resources for the duration of an active session.
The objective of the present study is not to reproduce the full packet-level behavior of the transport network, but to determine the minimum server-side delivery capacity required to satisfy differentiated service constraints for multiple traffic classes. Therefore, the OTT/IPTV architecture is reduced to a tractable resource model focused on the CDN edge server cluster. In this interpretation, each admitted request consumes a class-dependent amount of delivery capacity associated with a given service profile, bitrate, or subscription tier.
This abstraction is appropriate for first-order engineering dimensioning, where the key decision variable is the total amount of delivery resources that must be provisioned in order to satisfy class-specific blocking probability targets.
3.2. Modeling Assumptions
The analytical framework developed in this work is intended for first-order engineering dimensioning of OTT/IPTV delivery resources. The main bottleneck relevant to the considered planning task is assumed to be located at the CDN edge server cluster, while the intermediate transport network is treated as sufficiently provisioned and therefore not considered as a source of blocking. This assumption allows the analysis to focus on admission limitations caused by finite server-side delivery capacity. This assumption reflects the intended scope of the proposed framework, which focuses on first-order server-side capacity dimensioning. In operator-managed OTT/IPTV deployments, transport-network resources are commonly provisioned with a design margin relative to the CDN edge delivery platform; therefore, admission limitations may be dominated by finite server-side delivery capacity. The model should therefore be interpreted as a server-side dimensioning tool rather than as a complete end-to-end bottleneck model.
Request arrivals are modeled as a Poisson process, and service times are assumed to be exponentially distributed. These assumptions are standard in teletraffic modeling of loss systems and enable tractable analytical evaluation of occupancy distributions and class-specific blocking probabilities. Traffic classes differ in their resource demand, which reflects heterogeneous service characteristics such as bitrate, content profile, or subscription level.
The primary performance measure considered in the study is the class-specific blocking probability. A request is regarded as blocked when the available delivery capacity at the moment of arrival is insufficient to admit the requested session. Consequently, the proposed framework is not intended to capture packet-level dynamics, adaptive bitrate mechanisms, or transient transport impairments. Instead, it is designed as an analytical tool for capacity planning under differentiated service requirements.
3.3. Full-Availability Group Model
A full-availability group can be interpreted as a finite-capacity loss system in which incoming service requests occupy a specified number of allocation units. In the OTT/IPTV context considered here, the FAG model represents a pool of delivery resources available in the CDN edge server cluster. A request is admitted only if sufficient resources are available at the instant of arrival; otherwise, it is blocked and cleared from the system.
This abstraction is well suited to server-side OTT/IPTV dimensioning because each accepted session occupies a defined amount of delivery capacity for the duration of the session. As a result, the CDN edge server cluster can be represented as a shared finite-capacity resource pool in which heterogeneous service classes compete for common resources.
For a single traffic class with offered traffic intensity
A and system capacity
V, the occupancy distribution is given by
where
denotes the probability that
k allocation units are occupied in a system with capacity
V.
The corresponding blocking probability can be calculated using the Erlang B formula:
where
denotes the probability that an arriving request is rejected due to a lack of available capacity.
The single-class formulation provides the conceptual basis for the multi-class extension used in the remainder of the paper. In particular, it shows how finite-capacity server-side delivery resources can be linked to admission performance in terms of blocking probability.
3.4. Kaufman–Roberts Recursion for Multi-Class OTT Traffic
OTT/IPTV delivery systems typically serve heterogeneous traffic generated by services with different resource requirements. This heterogeneity may result from subscription differentiation, bitrate variation, or service-specific delivery profiles. Therefore, a multiservice extension of the loss model is required.
In the considered framework, the occupancy distribution for a resource group handling multiple traffic classes is determined using the Kaufman–Roberts recursion [
4,
13]:
where
denotes the offered traffic intensity of class
i,
is the resource demand of class
i, and
M is the number of traffic classes.
The blocking probability for class
i is obtained as the sum of occupancy states in which the remaining capacity is insufficient to accept a request of size
:
In the OTT/IPTV interpretation adopted in this paper, the parameter represents the amount of server-side delivery capacity required to handle a request from class i, whereas denotes the offered traffic intensity generated by that class.
This formulation is particularly suitable for multi-class OTT dimensioning because it directly links heterogeneous service demands to class-specific admission performance. As a result, it enables the estimation of the minimum delivery capacity required to satisfy differentiated blocking constraints for multiple subscription tiers within a unified analytical framework. For completeness and reproducibility, the procedural pseudocode used to implement the Kaufman–Roberts recursion is provided below.
From a computational perspective, the Kaufman–Roberts recursion has complexity per evaluation, where V is the system capacity (expressed in allocation units) and M is the number of traffic classes. For the considered scenarios (, ), this corresponds to approximately arithmetic operations per evaluation, which is negligible on modern hardware. The capacity search procedure evaluates the recursion for a sequence of candidate capacities, leading to a total complexity on the order of evaluations, where is the step size (equal to one allocation unit). The overall computational cost therefore remains low. It should also be noted that the method does not rely on iterative numerical convergence. Each evaluation of the recursion yields an exact occupancy distribution for the given parameters, and the capacity search procedure is deterministic.
3.5. Validation Procedure
Before applying the analytical framework to the dimensioning scenarios, the model was validated against discrete-event simulation written in C# version 11. The objective of this step was to verify whether the proposed analytical formulation provides reliable estimates of class-specific blocking probabilities over the considered operating range.
Each run spanned simulated time units; the first 5000 were excluded to allow the system to reach steady state. For each operating point, twelve independent runs were performed, and the results were reported together with 95% confidence intervals computed using Student’s t-distribution. The reference validation scenario consisted of a full-availability group with a total capacity of 360 Mbit/s and six traffic classes requesting 4 Mbit/s, 8 Mbit/s, 12 Mbit/s, 18 Mbit/s, 24 Mbit/s, and 30 Mbit/s, respectively. An equal traffic load ratio was assumed for all classes, i.e., the product of offered traffic intensity and resource demand was identical across classes. In the prioritised validation case, class and had the highest priority, followed by and with medium priority followed by and with low priority.
The agreement between the analytical and simulation results was assessed by comparing the blocking probabilities obtained for the individual traffic classes presented in the figures in the Results section. Close agreement between the two approaches confirms that the adopted analytical abstraction is sufficiently accurate for subsequent engineering dimensioning. Moreover, it should be noted that the underlying analytical formulation, based on the Kaufman–Roberts recursion, provides an exact solution for the considered class of loss network models, and its accuracy does not depend on the scale of the system but on the validity of the modeling assumptions (e.g., traffic independence and stationarity). Therefore, the purpose of the simulation-based validation is not to establish correctness, but to confirm that these assumptions provide an adequate representation of the system under study. This supports the application of the model to larger and more heterogeneous dimensioning scenarios, provided that the underlying stochastic assumptions (e.g., traffic independence and stationarity) remain valid.
Three resource provisioning strategies were considered in the dimensioning study. These strategies represent distinct operator-side resource management policies, ranging from strict resource separation to full pooling with service differentiation.
In the first strategy, dedicated servers were assigned to the free, standard, and premium subscription tiers. The free tier included traffic classes and with resource demands of 6 Mbit/s and 8 Mbit/s, respectively. The standard tier included classes and , with resource demands of 10 Mbit/s and 12 Mbit/s. The premium tier included classes and with resource demands of 14 Mbit/s and 16 Mbit/s. An equal traffic load ratio was assumed for all classes.
In the second strategy, all traffic classes were aggregated into a single shared non-prioritised resource pool. In the third strategy, all traffic classes were served by a single shared prioritised server, where premium traffic had the highest priority, standard traffic had intermediate priority, and free-tier traffic had the lowest priority.
The blocking probability targets of 1%, 0.1%, and 0.001% were adopted to represent progressively stricter service-continuity requirements for free, standard, and premium subscription tiers, respectively. In this study, these values are used as engineering design targets corresponding to differentiated service expectations in operator-managed OTT/IPTV environments, consistent with commonly adopted teletraffic engineering practice.
The selected throughput values represent differentiated OTT/IPTV service profiles corresponding to fixed bitrate service classes (e.g., associated with different video quality levels) and are consistent with operator-side engineering practice. While not tied to a single dataset, these values reflect configurations observed in production environments, including representative OTT/IPTV deployment scenarios discussed in the literature [
1,
11].
The traffic range considered in the numerical study extended from 2000 Mbit/s to 6000 Mbit/s. Assuming equal representation of the six traffic classes and the considered aggregate traffic range, this corresponds to approximately 180 to 540 simultaneously active STBs. Assuming an activity factor of 0.25, this translates into a total deployed population of approximately 725 to 2180 STBs, which reflects a moderate-scale IPTV deployment scenario.
Together, these three strategies span a meaningful engineering design space, ranging from strict isolation to full pooling and controlled differentiation.
3.6. Capacity Search Procedure
The minimum required server capacity for each provisioning strategy was determined using an incremental search procedure. For a given set of offered traffic intensities
and blocking probability targets
, the procedure identifies the smallest integer capacity
expressed in allocation units such that all class-specific blocking constraints are simultaneously satisfied, i.e.,
where
denotes the blocking probability of class
i in a system with capacity
V, computed using the Kaufman–Roberts recursion (Equation (
3)) for the non-prioritised configurations or using the corresponding priority-aware formulation for the prioritised configuration.
The search is initialized at a lower bound and proceeds by evaluating consecutive feasible capacity values until all class-specific blocking constraints are satisfied. This makes it possible to determine the minimum delivery capacity required under the assumed traffic mix and service requirements.
Equivalently, expressing all resource demands and the candidate capacity in units of yields integer-valued demands and integer capacity , as required by the Kaufman–Roberts recursion (Appendix). For the considered scenario, Mbit/s, so and a candidate capacity of Mbit/s corresponds to integer states. This normalisation does not affect the computed blocking probabilities, since multiplying all resource quantities by the same constant leaves the occupancy distribution invariant.
From an engineering perspective, this procedure provides a direct and reproducible decision rule for infrastructure planning, since it identifies the minimum server-side delivery capacity required to satisfy all class-specific service constraints under the assumed traffic mix.
The monotonicity of the class-specific blocking probabilities
as non-increasing functions of the system capacity
V is a standard property of multiservice loss systems (see [
17], Ch. 7). This property follows from the stochastic ordering of occupancy distributions, as increasing
V enlarges the admissible state space and reduces the probability that the available capacity is insufficient for a given class. In the prioritised configuration (
Section 4.4), the blocking probabilities are constructed from nested full-availability subsystems, each of which is itself a standard FAG. Therefore, the monotonicity property also holds for the prioritised blocking probabilities. As a consequence, the incremental capacity search procedure is guaranteed to terminate at the minimal capacity
satisfying all class-specific constraints.
Figure 1 summarises the overall workflow of the proposed analytical capacity-dimensioning framework.
For readability, the main text presents only the analytical formulations and their engineering interpretation. The procedural pseudocode used to implement the recursion-based computations is collected in this appendix. Algorithm 1 gives the Kaufman–Roberts recursion for a non-prioritised full-availability group, whereas Algorithm 2 shows the nested procedure used to evaluate blocking probabilities in the prioritised configuration.
| Algorithm 1 Kaufman–Roberts recursion for a multi-class full-availability group. |
- Require:
System capacity V, number of classes M, offered traffic vector , resource-demand vector - Ensure:
Occupancy probabilities for and class-specific blocking probabilities - 1:
Initialize - 2:
for to V do - 3:
- 4:
for to M do - 5:
if then - 6:
- 7:
end if - 8:
end for - 9:
- 10:
end for - 11:
Compute normalization constant: - 12:
for to V do - 13:
- 14:
end for - 15:
for to M do - 16:
- 17:
end for
|
| Algorithm 2 Priority-aware blocking computation using nested full-availability groups. |
- Require:
System capacity V, number of classes M, offered traffic vector , resource-demand vector ordered from highest to lowest priority - Ensure:
Priority-aware blocking probabilities for all classes - 1:
for to M do - 2:
Consider subsystem containing classes - 3:
Compute occupancy probabilities using Algorithm 1 restricted to classes - 4:
for to m do - 5:
Compute blocking probability in subsystem m: - 6:
end for - 7:
end for - 8:
for to M do - 9:
if then - 10:
- 11:
else - 12:
Compute priority-aware blocking: - 13:
end if - 14:
end for
|
4. Results
4.1. Validation Results
Figure 2 compares the analytical results obtained with the proposed framework and the corresponding simulation results. Close agreement between the two approaches can be observed over the entire range of offered traffic. The simulation confidence interval half–widths are presented in
Table 1 and
Table 2. This confirms the correctness of the adopted analytical formulation and supports its use as the primary evaluation tool in the subsequent dimensioning scenarios.
4.2. Dedicated-Server Configuration
The first dimensioning scenario considers three separate servers assigned to the free, standard, and premium subscription tiers. The objective is to determine the minimum required capacity of each server such that the assumed blocking probability constraints are satisfied over the considered traffic range.
The obtained results show that the required capacities are equal to 2180 Mbit/s for the free subscription, 2406 Mbit/s for the standard subscription, and 2744 Mbit/s for the premium subscription. The corresponding summary is presented in
Table 3, and the blocking-probability profiles as a function of offered traffic are illustrated in
Figure 3.
It should be noted that the total capacity reported for the dedicated-server configuration is obtained as the sum of three independently dimensioned subsystems corresponding to the individual subscription tiers. Each subsystem is modeled as a separate full-availability group and dimensioned independently to satisfy its own class-specific blocking constraints. Therefore, no assumption of capacity additivity across different provisioning strategies is made.
These results confirm that stricter service guarantees require progressively larger dedicated capacities. In particular, the premium tier, characterized by the most restrictive blocking target, requires the largest amount of reserved server-side resources. The minimum required capacities and the corresponding blocking targets are summarised in
Table 3; the total capacity required by the dedicated-server architecture is 7330 Mbit/s.
4.3. Shared-Server Configuration
The second scenario considers a single shared non-prioritised server handling all traffic classes. In this case, the objective is to determine the minimum shared capacity required to satisfy the same class-specific blocking probability constraints as in the dedicated-server architecture.
The analytical results for this configuration are presented in
Table 4. They show that the minimum server capacity required to satisfy all blocking probability constraints is equal to 7050 Mbit/s.
Compared with the dedicated-server architecture, whose total required capacity is 7330 Mbit/s, the shared-server configuration reduces the capacity requirement by 280 Mbit/s, which corresponds to approximately 3.82%. This reduction is a direct consequence of statistical multiplexing, which improves the utilization of shared resources while preserving the required class-specific blocking constraints. As visible in
Table 4, the blocking probabilities of all six classes are closely comparable at each traffic level, reflecting the non-prioritised nature of the shared resource pool in which all classes experience the same level of congestion.
The extremely low blocking probabilities observed at low offered traffic levels are a direct consequence of the system operating far below its capacity. In this regime, the utilization is low and sufficient resources are almost always available, so blocking can only occur due to very rare stochastic fluctuations. As a result, the blocking probability decreases exponentially and reaches values that are practically negligible (e.g., below
). This behavior is characteristic of loss systems and reflects the rapidly decaying tail of the occupancy distribution when the offered traffic is significantly smaller than the system capacity. Most of such extreme values were omitted in
Table 4 as they do not impact the resource dimensioning decisions.
4.4. Prioritized Shared-Server Configuration
The third scenario considers a single shared server with differentiated access priorities. In the OTT/IPTV setting considered, premium traffic classes are assigned the highest priority, standard traffic classes have intermediate priority, and free-tier traffic classes have the lowest priority. As a result, higher-priority traffic is protected against congestion at the expense of increased blocking experienced by lower-priority traffic.
The adopted priority mechanism is based on the analytical procedure proposed in [
14,
15]. Although the cited studies considered radio-access applications, the priority mechanism itself is formulated at the level of a multi-service full-availability group. Therefore, its applicability is not restricted to WCDMA or LTE systems. The same mathematical assumptions are used in the present work: a finite homogeneous resource pool, Poisson arrivals, exponentially distributed service times, class-dependent integer resource demands, and loss admission.
The priority mechanism is interpreted at the admission level. Higher-priority traffic is protected by evaluating its blocking probability in nested subsystems that exclude lower-priority classes, whereas lower-priority traffic is evaluated under the cumulative load of higher-priority classes. Thus, lower-priority traffic affects neither the admission conditions nor the blocking probabilities of higher-priority classes. For a system with capacity
V, class-specific resource demands
, and offered traffic intensities
, the blocking probability for class
k in a prioritised system can be determined by the following formula [
14,
15]:
where
denotes the blocking probability of class
k in the prioritised system,
denotes the blocking probability of class
i computed for the
k-class subsystem
, and
is the corresponding value for the
-class subsystem
. The formula follows from a nested decomposition of the FAG state space. Classes are ordered from the highest to the lowest priority. For class
k, the subsystem
represents the resource competition visible to that class. The difference
captures the additional loss contribution introduced when class
k is added to the subsystem. Consequently, Equation (
6) is not a heuristic adaptation from wireless systems, but a direct application of the priority-aware FAG model to the OTT/IPTV server-side resource pool. Internal fragmentation is not present in the considered model. The resource pool is homogeneous and consists of indistinguishable allocation units. A request requiring
units is admitted whenever at least
units are available, regardless of their position. Hence, unlike systems with contiguity constraints, such as spectrum-slot, memory, or wavelength allocation systems, the feasibility of admission depends only on the total occupied capacity and not on the arrangement of free units.
The procedural pseudocode corresponding to the priority-blocking computation is provided above. Classes are assumed to be ordered from highest () to lowest () priority, so that the Kaufman–Roberts recursion is applied iteratively to nested subsystems of increasing size.
The results obtained for the prioritised shared-server configuration are presented in
Table 5. An analysis of the results shows that the prioritisation mechanism strongly suppresses the blocking probability of higher-priority classes, while shifting a larger share of blocking events to lower-priority traffic. At the highest traffic load considered (
Mbit/s), the blocking probabilities of the premium traffic classes remain extremely small (of the order of
), whereas the corresponding values for the free-tier classes increase to the order of
—comparable to the dedicated-server case.
This behavior confirms that the prioritised model effectively protects premium traffic against congestion while preserving explicit service differentiation within a common resource pool.
The minimum required capacity of the prioritised shared-server configuration was equal to 6418 Mbit/s. Compared with the dedicated-server architecture, the prioritised shared-server configuration reduces the total required capacity by approximately %, while simultaneously enabling explicit service differentiation across subscription tiers.
The extremely low blocking probabilities observed at low offered traffic levels are a direct consequence of the system operating far below its capacity: blocking can only occur due to very rare stochastic fluctuations, so the occupancy distribution decays exponentially into values that are numerically negligible. Prioritization further accentuates this effect by allowing higher-priority traffic to effectively ignore resource occupancy caused by lower-priority classes.
A cross-strategy comparison of the blocking probabilities at the peak offered traffic of
Mbit/s is provided in
Figure 4. The grouped bar chart illustrates the redistribution of blocking load across classes: in the dedicated and shared configurations all classes remain close to their respective thresholds, whereas in the prioritised configuration the standard- and premium-tier classes are driven many orders of magnitude below their targets at the direct expense of the free-tier classes, which absorb the excess blocking. This redistribution is the mechanism by which a tighter common capacity can still satisfy all GoS constraints simultaneously.
5. Discussion
The obtained results show that the resource provisioning strategy has a direct impact on both the total required capacity and the achievable level of service differentiation. The capacity requirements for all three strategies are summarised in
Table 6.
It should be noted that the dedicated-server configuration considered in this study is not intended as a worst-case construct, but rather as a practical and commonly used reference point in operator deployments, where service differentiation is often realised through physically or logically separated resource pools. The reported capacity savings should therefore be interpreted relative to this baseline, reflecting the gains achievable when moving from strict isolation toward shared and prioritised resource allocation.
At the same time, the presented study should be interpreted within the limitations of the adopted model. The analysis is based on classical Erlang traffic assumptions, i.e., Poisson arrivals and exponentially distributed service times. These assumptions are used here at the session-admission level rather than at the packet or segment level. They provide a tractable approximation when the aggregate demand is generated by a sufficiently large user population and when the planning objective is to estimate admission-level blocking probabilities.
OTT/IPTV traffic may nevertheless exhibit non-stationary behaviour, including time-of-day variation, event-driven bursts, and correlated user activity. If such effects are stronger than assumed in the model, the blocking probabilities obtained from stationary offered traffic may underestimate the true blocking risk. In practical planning, this can be mitigated by applying the model to busy-hour or peak-hour equivalent traffic and by calibrating the offered load using measured operator-side traces.
Adaptive bitrate mechanisms constitute another source of model abstraction. In practice, the bitrate of an active stream may change over time in response to network and client conditions. In the proposed framework, this variability is represented through class-dependent resource demands , which may be interpreted as average, design, or conservative upper bitrate values associated with a given service profile. A more detailed representation of adaptive bitrate dynamics would require an extended model with variable resource demands, for example based on BPP-type or state-dependent traffic formulations.
Finally, the model assumes that the main bottleneck relevant to the planning task is located at the CDN edge server cluster, while the intermediate transport network is sufficiently provisioned. This assumption is appropriate for server-side dimensioning. In scenarios where multiple bottlenecks are relevant, the approach could be extended by applying the FAG formulation to consecutive resource pools or by developing cascaded/multi-resource loss models.
It is also important to distinguish the present analytical dimensioning framework from QoE-driven online resource allocation approaches. In contrast to real-time scheduling or adaptive streaming control, where computational latency directly affects system performance and may require latency-aware evaluation metrics, e.g., [
18], the proposed method is intended for offline infrastructure planning. As a result, computational complexity does not influence the dimensioning outcome, provided that it remains tractable.
The dedicated-server architecture offers the strongest logical separation between subscription tiers, but it also requires the largest total capacity because resources are reserved independently for each traffic group. Its capacity requirement lies consistently above both shared-server alternatives across the entire traffic range considered.
The shared non-prioritised configuration improves capacity efficiency by exploiting statistical multiplexing across all traffic classes. The capacity saving relative to the dedicated-server baseline is converging to approximately 3.8% at Mbit/s.
The prioritised shared-server strategy achieves the largest capacity reduction, with savings ranging from approximately 15% at low traffic to 12% at the highest considered load. From the operator perspective, this result demonstrates that traffic prioritisation not only provides service differentiation but also yields meaningful infrastructure savings compared with both dedicated and non-prioritised shared provisioning.
The per-tier behavior underlying these capacity savings can be inferred from the blocking-probability profiles presented in the Results section. For the free tier, the prioritised configuration produces higher blocking than both the dedicated and shared alternatives at the same total traffic level. This is the direct consequence of the priority mechanism, which shifts congestion away from higher-priority classes toward lower-priority ones. For the standard and premium tiers, the prioritised strategy consistently achieves lower blocking than the dedicated configuration, confirming that the gain in capacity efficiency is accompanied by improved congestion protection for higher-priority traffic. The shared non-prioritised configuration occupies an intermediate position for the free tier and shows slightly higher blocking than the dedicated configuration for premium traffic, reflecting the fact that the shared pool must satisfy all classes simultaneously under a common resource constraint.
It should also be noted that the introduction of priorities leads to a significant redistribution of blocking across traffic classes. In particular, the blocking probabilities for the highest-priority classes are substantially lower than those obtained in the non-prioritised configuration, whereas the values for lower-priority classes increase accordingly. This behavior is consistent with the conservation of traffic carried by the system and confirms that prioritisation acts as an internal load-balancing mechanism rather than a global capacity enhancement.