4.2. Overall Results
Table 2 summarizes the overall performance on the three datasets. The proposed model achieves SROCC values of 0.91537, 0.9101, and 0.9778 on SJTU-PCQA, WPC, and ICIP2020, respectively. These results indicate strong ranking consistency across datasets with different content compositions and distortion distributions. The corresponding PLCC values further show that the fused representation supports a stable mapping to MOS values.
To compare the proposed model with representative FR-PCQA baselines,
Table 3 reports the results of a broad set of competing methods on the three datasets. The baselines include projection domain metrics such as PSNR, SSIM, MS-SSIM, and VIF, point domain or hybrid metrics such as MSE-p2po, MSE-p2pl, PCQM, GraphSIM, and MS-GraphSIM, and learning based or perception guided baselines such as FRSVR, PHM, and PW-SCQA.
All competing methods in
Table 3 were reproduced using the same content-level splits as the proposed model. For methods whose raw quality scores are not expressed on the MOS scale, a four-parameter logistic function was fitted using only the training contents within each outer fold and then applied to the corresponding test contents before PLCC and RMSE were computed. For these methods, SROCC was calculated directly from the raw scores.
Table 3 shows that the proposed model achieves the highest PLCC and SROCC values and the lowest RMSE on the three evaluated datasets. Compared with representative point domain metrics such as PCQM, GraphSIM, and MS-GraphSIM, the gains are particularly clear on WPC and SJTU-PCQA. The proposed model also performs favorably compared with the learning based baseline FRSVR. These results suggest that the proposed FR-PCQA method can provide a more complete description of the distortions.
The best performance is achieved on the ICIP2020 dataset, where distortions are dominated by compression and are relatively concentrated. SJTU-PCQA and WPC are more challenging because they contain broader distortion diversity and larger perceptual variability. Overall, the proposed dual branch model maintains stable performance across all three datasets, indicating robustness under heterogeneous distortion conditions.
To examine generalization under changes in data distribution, we conduct six directed database transfer experiments among SJTU-PCQA, WPC, and ICIP2020. For each experiment, the model is trained on one database and directly evaluated on another. Missing values are completed using statistics estimated from the training database, and no score mapping is estimated from the test database. SROCC, PLCC, and KRCC are reported because the databases use different MOS scales.
As shown in
Table 4, transfer performance varies with the training and test databases. Training on SJTU-PCQA and WPC yields SROCC values of 0.8982 and 0.8740 on ICIP2020, respectively, while training on WPC and testing on SJTU-PCQA yields an SROCC of 0.8195. The lowest SROCC of 0.5816 is obtained when the model is trained on ICIP2020 and tested on WPC. These results demonstrate the ability of the proposed model to transfer across databases, while also showing that its performance is influenced by differences between the training and test data distributions.
Robustness under controlled external perturbations is further examined using 18 samples from SJTU-PCQA, with two samples selected from approximately the lower and upper MOS tertiles of each source content. Each sample is evaluated using a model trained without its source content. Joint rotations of 15 and 30 degrees about the axis are applied to both the reference and distorted point clouds to examine pose stability. Geometry noise with standard deviations of 0.05% and 0.10% of the reference bounding box diagonal and random point removal ratios of 5% and 10% are applied to the distorted point cloud. A fixed random seed of 2026 is used.
Table 5 shows that joint rotation largely preserves the prediction ranking at 15 degrees, while the lower SROCC at 30 degrees indicates increased sensitivity to larger changes in orientation. Geometry noise produces a clear decrease in the predicted quality, and the stronger noise level results in a lower prediction than the weaker level for all 18 samples. Random point removal maintains high rank correlations, but its effect on the absolute predictions is not consistently directional. These results indicate that the response of the proposed model depends on both the type and intensity of the external perturbation.
4.6. Ablation Study
To quantify the contribution of each internal module and the final fusion strategy, seven ablation settings are evaluated: DISTS only, projection only, global structure only, DISTS + projection, DISTS + projection + global structure, DISTS + projection + PCQM, and full model (GBRT). Although the architecture contains two main branches, the ablations are conducted at the module level to identify the source of the complementary gain. A dedicated pairwise comparison is also reported to isolate the internal complementarity of the 2D and 3D branches. The overall ablation trend is shown in
Figure 2.
Table 9 directly evaluates the two branch level complementary hypotheses. In the 2D branch, DISTS + projection consistently outperforms DISTS only on all three datasets and exceeds projection only on SJTU-PCQA and WPC. On ICIP2020, projection only gives the highest SROCC and PLCC within the 2D branch, whereas DISTS + projection gives a higher KRCC, increasing it from 0.8467 to 0.8562, and reduces RMSE from 0.4899 to 0.4637. This pattern indicates that combining perceptual similarity with occupancy mismatch, geometric deviation, and chromatic discrepancy provides more balanced performance across the evaluation criteria.
In the 3D branch, the complementarity is stronger and more consistent. PCQM + global structure outperforms both PCQM only and global structure only on SJTU-PCQA, WPC, and ICIP2020. Compared with PCQM only, the SROCC gains are +0.0396, +0.0133, and +0.0048 on the three datasets, respectively, and PLCC also improves in all cases. This observation is consistent with the role of the global descriptor.
On SJTU-PCQA, the SROCC values of DISTS only, projection only, and global structure only are 0.7286, 0.7826, and 0.8459, respectively. These values indicate that global structural statistics alone are more informative than either 2D component on this dataset. When DISTS and projection features are combined, SROCC increases to 0.8324. Adding the 3D global structure module further increases SROCC to 0.9144, which is already close to the full model result. This behavior is consistent with the distortion composition of SJTU-PCQA, where downscaling and geometry noise strongly affect point count change, scale variation, spatial redistribution, and density variation. Accordingly, bounding box scale, global scale, radial distribution, volumetric density, and covariance spectrum descriptors provide important evidence that is not sufficiently captured by local PCQM responses. Adding PCQM only increases SROCC from 0.9144 to 0.91537, indicating that the dominant gain on this dataset comes from global structural compensation, while local geometry and color distortion cues mainly refine PLCC and RMSE.
On WPC, the SROCC values of DISTS only, projection only, and global structure only are 0.7617, 0.6299, and 0.5875, respectively. These values indicate that no single module can adequately describe the complex distortions in this dataset. Starting from the 2D combination DISTS + projection, adding either the global structure module or the PCQM module increases SROCC to 0.8832 and 0.8937, respectively, while the full model further reaches 0.9101. This pattern suggests that WPC contains both point count change, scale variation, spatial redistribution, and density variation described by global structural statistics, as well as local geometry and color distortion described by PCQM evidence. Thus, both components of the 3D branch are necessary.
On ICIP2020, the projection only configuration already reaches an SROCC of 0.9498, which is higher than the values of DISTS only (0.8921) and global structure only (0.8668). This indicates that occupancy mismatch, geometric deviation, and chromatic discrepancy in the 2D projection domain are highly sensitive to compression distortion. The full model further improves SROCC to 0.9778, showing that local geometry and color distortion, point count change, scale variation, spatial redistribution, and density variation still provide stable gains even when the 2D explicit error module is strong.
In addition to the module-based ablation experiments, we further evaluate the performance of the six groups of features in the proposed 2D branch and the eight groups of features in the 3D branch. In the experiment, the SJTU-PCQA dataset is employed since it contains geometry, color, downsampling, octree compression, and mixed distortions. To evaluate the contribution of a specific feature group, we remove it from the full feature set while maintaining all other settings. The contribution is then measured as the difference between the SROCC value of the complete feature set and that of the set without the target group. Therefore,
is defined as
, where a larger positive value indicates a greater contribution. The specific results for each feature group are shown in
Table 10.
As shown in
Table 10, all the
values are positive, indicating that all the proposed feature groups contribute positively to the final quality prediction. The maximum and minimum values are 0.0185 and 0.0011, respectively. Specially, for the projection module, the
values for Occupancy, Depth Fidelity, Gradient Domain Fidelity, Normal Consistency, Color Fidelity, and Luma Fidelity are 0.0027, 0.0031, 0.0026, 0.0035, 0.0011, and 0.0018, respectively. For global structure module, the
values for Point Count Ratio, Centroid Offset, Bounding Box Axis Scale Difference, Spatial Dispersion Difference, Covariance Spectrum Difference, Radial Statistics Difference, Global Scale Difference, and Volume Density Difference are 0.0185, 0.0017, 0.0013, 0.0057, 0.0044, 0.0021, 0.0013, and 0.0012, respectively.
4.8. Small Sample Robustness
To evaluate sensitivity to the amount of available training content, we conduct repeated content sampling under content independent outer fold protocols. SJTU-PCQA and ICIP2020 use one content out protocols with 9 and 6 outer folds, respectively, while WPC uses 10 content independent outer folds, each containing two test contents. For each outer fold, the test contents remain fixed, and the specified number of reference contents is sampled without replacement from the remaining training pool. All distorted point clouds derived from a selected reference content are retained together in the training set. The numbers of training contents are for SJTU-PCQA, for WPC, and for ICIP2020. For each reduced setting, the sampling procedure is repeated 20, 12, and 20 times per outer fold, respectively. The full setting is evaluated once because it contains the complete training pool of the corresponding outer fold. A base seed of 2026 is used, and the seed for repetition r in outer fold k with m training contents is .
The SROCC values obtained from repeated training subsets are first averaged within each outer test fold. The resulting 9, 10, and 6 outer fold means are used to calculate the overall mean and sample standard deviation. The 95% confidence intervals (CIs) are estimated using 20,000 iterations of a hierarchical percentile bootstrap that resamples the outer folds and the repeated training subsets within each selected fold.
Table 13 reports the mean, standard deviation, and confidence interval, and
Figure 3 shows the corresponding means and 95% CIs.
For the paired comparisons, each fold mean obtained with a reduced setting is compared with the full setting evaluated on the same outer test fold. We report
, its paired hierarchical bootstrap 95% CI, and the result of a two-sided exact permutation test based on sign changes in the fold level differences. The resulting
p values are adjusted within each dataset using the Holm procedure, as reported in
Table 14.
The results show that the difference from the full setting generally decreases as the number of training contents increases. Compared with the corresponding full setting, the use of two training contents on SJTU-PCQA and four training contents on WPC decreases SROCC by 0.0397 and 0.0271, respectively. Both differences are statistically significant after Holm correction. The differences for six versus eight training contents on SJTU-PCQA and for 8, 12, or 16 versus 18 training contents on WPC are not statistically significant. For ICIP2020, the mean SROCC differences range from 0.0036 to 0.0116, and none of the paired comparisons is statistically significant after Holm correction. The standard deviations and confidence intervals quantify the variation across outer test folds under the evaluated training settings.