Figure 1.
Multi-head neural network architecture. A shared backbone (dashed box) compresses 29 input features into a 128-dimensional device-state representation. Four task-specific prediction heads, with depths calibrated to target complexity, branch from this shared representation to produce scalar estimates of , , FF, and PCE. Total model parameters: ≈83,000.
Figure 1.
Multi-head neural network architecture. A shared backbone (dashed box) compresses 29 input features into a 128-dimensional device-state representation. Four task-specific prediction heads, with depths calibrated to target complexity, branch from this shared representation to produce scalar estimates of , , FF, and PCE. Total model parameters: ≈83,000.
Figure 2.
Experimental design overview.
Figure 2.
Experimental design overview.
Figure 3.
Feature engineering ablation. Horizontal bars show for each target under six feature configurations. All configurations achieve , demonstrating that the multi-task architecture is the primary driver of prediction accuracy, with physics-guided features providing marginal gains in inter-parameter consistency.
Figure 3.
Feature engineering ablation. Horizontal bars show for each target under six feature configurations. All configurations achieve , demonstrating that the multi-task architecture is the primary driver of prediction accuracy, with physics-guided features providing marginal gains in inter-parameter consistency.
Figure 4.
Loss weighting sensitivity heatmap. per target across six weight configurations. All cells exceed 0.994, confirming that prediction accuracy is invariant to the choice of loss weights within the tested range. On the vertical axis, “↓” denotes a down-weighted task.
Figure 4.
Loss weighting sensitivity heatmap. per target across six weight configurations. All cells exceed 0.994, confirming that prediction accuracy is invariant to the choice of loss weights within the tested range. On the vertical axis, “↓” denotes a down-weighted task.
Figure 5.
Physical consistency error by loss weighting. Uniform weighting achieves the lowest consistency MAE (0.222), while increasing FF emphasis progressively degrades consistency. Error bars indicate cross-fold standard deviation. Error bars indicate cross-fold standard deviation ().
Figure 5.
Physical consistency error by loss weighting. Uniform weighting achieves the lowest consistency MAE (0.222), while increasing FF emphasis progressively degrades consistency. Error bars indicate cross-fold standard deviation. Error bars indicate cross-fold standard deviation ().
Figure 6.
Per-fold for single-task (red) and multi-task (blue) models across all four targets. Single-task FF ranges from 0.321 to 0.907, while multi-task maintains in all folds. The 233× reduction in cross-fold standard deviation demonstrates qualitative transformation from unreliable to stable prediction.
Figure 6.
Per-fold for single-task (red) and multi-task (blue) models across all four targets. Single-task FF ranges from 0.321 to 0.907, while multi-task maintains in all folds. The 233× reduction in cross-fold standard deviation demonstrates qualitative transformation from unreliable to stable prediction.
Figure 7.
Distribution of signed consistency errors. Multi-task predictions (blue) are tightly concentrated near zero, while single-task predictions (red) exhibit heavy tails. Dashed lines indicate the ±2 PCE-unit implausibility threshold. 36.5% of single-task predictions exceed this threshold versus 0.014% for multi-task.
Figure 7.
Distribution of signed consistency errors. Multi-task predictions (blue) are tightly concentrated near zero, while single-task predictions (red) exhibit heavy tails. Dashed lines indicate the ±2 PCE-unit implausibility threshold. 36.5% of single-task predictions exceed this threshold versus 0.014% for multi-task.
Figure 8.
Model comparison across all four targets. The MH-NN achieves the highest on every target with the tightest cross-fold variability. CatBoost is the strongest tree-based competitor, while XGBoost shows the weakest performance, particularly on FF. Error bars indicate cross-fold standard deviation ().
Figure 8.
Model comparison across all four targets. The MH-NN achieves the highest on every target with the tightest cross-fold variability. CatBoost is the strongest tree-based competitor, while XGBoost shows the weakest performance, particularly on FF. Error bars indicate cross-fold standard deviation ().
Figure 9.
Per-composition (MT − ST) heatmap across 12 perovskite compositions and four targets. Most differences are negligible (), confirming that the multi-task advantage arises from fold-level stabilization rather than composition-specific accuracy gains.
Figure 9.
Per-composition (MT − ST) heatmap across 12 perovskite compositions and four targets. Most differences are negligible (), confirming that the multi-task advantage arises from fold-level stabilization rather than composition-specific accuracy gains.
Figure 10.
Permutation importance (50 repetitions, 95% CI) for the top features across all four targets. The I/Br ratio dominates prediction (0.631), FA fraction dominates (0.557) and PCE (0.523), and multiple ETL categories appear as critical features for , FF, and PCE.
Figure 10.
Permutation importance (50 repetitions, 95% CI) for the top features across all four targets. The I/Br ratio dominates prediction (0.631), FA fraction dominates (0.557) and PCE (0.523), and multiple ETL categories appear as critical features for , FF, and PCE.
Figure 11.
SHAP beeswarm summary for all four targets. Red indicates high feature values, blue indicates low. The analysis confirms the halide-mediated – trade-off (I/Br shows opposing signs across targets), the thickness– Beer–Lambert relationship, and ETL-dominated FF and PCE control.
Figure 11.
SHAP beeswarm summary for all four targets. Red indicates high feature values, blue indicates low. The analysis confirms the halide-mediated – trade-off (I/Br shows opposing signs across targets), the thickness– Beer–Lambert relationship, and ETL-dominated FF and PCE control.
Figure 12.
SHAP dependence plots for selected feature–target pairs. Non-linear relationships and interaction effects are visible, including saturation behavior in the I/Br– relationship and FA-dependent stratification in the I/Br–PCE relationship.
Figure 12.
SHAP dependence plots for selected feature–target pairs. Non-linear relationships and interaction effects are visible, including saturation behavior in the I/Br– relationship and FA-dependent stratification in the I/Br–PCE relationship.
Figure 13.
Hyperparameter sensitivity across four studies: backbone width, backbone depth, dropout rate, and batch size. for all four targets remains above 0.99 across most configurations, with FF consistently the most sensitive target. The chosen configuration (marked) lies in the plateau region for all studies.
Figure 13.
Hyperparameter sensitivity across four studies: backbone width, backbone depth, dropout rate, and batch size. for all four targets remains above 0.99 across most configurations, with FF consistently the most sensitive target. The chosen configuration (marked) lies in the plateau region for all studies.
Figure 14.
Dataset-size scaling comparison between MH-NN and CatBoost. Left: per-target as a function of training set size. Right: consistency MAE as a function of training set size. CatBoost dominates at ; MH-NN overtakes on at with consistency converging at the full dataset.
Figure 14.
Dataset-size scaling comparison between MH-NN and CatBoost. Left: per-target as a function of training set size. Right: consistency MAE as a function of training set size. CatBoost dominates at ; MH-NN overtakes on at with consistency converging at the full dataset.
Figure 15.
Sensitivity of the implausible-prediction rate to the consistency threshold. The multi-task model reaches zero violations above 2.0 PCE units, while the single-task model retains substantial violation rates across all thresholds.
Figure 15.
Sensitivity of the implausible-prediction rate to the consistency threshold. The multi-task model reaches zero violations above 2.0 PCE units, while the single-task model retains substantial violation rates across all thresholds.
Figure 16.
Noise injection robustness study with multiplicative Gaussian noise (– relative measurement error). (Left) per-target degradation. The MT (green) and MT PCE (pink) curves nearly coincide and appear as a single curve. (Right) consistency MAE degradation. The multi-task model degrades monotonically and maintains better consistency at than the single-task model on clean data.
Figure 16.
Noise injection robustness study with multiplicative Gaussian noise (– relative measurement error). (Left) per-target degradation. The MT (green) and MT PCE (pink) curves nearly coincide and appear as a single curve. (Right) consistency MAE degradation. The multi-task model degrades monotonically and maintains better consistency at than the single-task model on clean data.
Table 1.
Comparison with prior ML approaches for PSC parameter prediction.
Table 1.
Comparison with prior ML approaches for PSC parameter prediction.
| Reference | Method | Data Source | Samples | Compositions | Targets | Consistency |
|---|
| Novoselov [9] | CatBoost | SCAPS-1D | 7182 | 12 | PCE only | No |
| Reza [10] | Rand. Forest | SCAPS-1D | 1000 | 1 | 4 (indep.) | No |
| Li [11] | Extra Trees | Experim. | 847 | Multiple | Per-output | No |
| Saidani [12] | NN + SHAP | SCAPS-1D | — | 1 | Per-output | No |
| This work | MH-NN (MTL) | SCAPS-1D | 7176 | 12 | 4 (joint) | Yes |
Table 2.
Target variable distributions ().
Table 2.
Target variable distributions ().
| Target | Min | Max | Mean | Std | Median | N |
|---|
| (V) | 0.934 | 2.005 | 1.309 | 0.228 | 1.243 | 7176 |
| (mA/cm2) | 2.488 | 28.500 | 17.670 | 7.644 | 20.544 | 7176 |
| FF (%) | 14.759 | 90.876 | 73.692 | 14.658 | 81.685 | 7176 |
| PCE (%) | 2.154 | 30.988 | 16.517 | 7.621 | 17.780 | 7176 |
Table 3.
Physics-guided engineered features.
Table 3.
Physics-guided engineered features.
| Feature | Formula | Physical Justification |
|---|
| Cs × MA | Cs · MA | A-site cation interaction: non-linear mixing effects on lattice stability and bandgap |
| I × Br | I · Br | Halide interaction: mixed-halide effects on absorption edge and defect chemistry |
| Cs/FA | Cs/(FA + ) | Cation ratio: governs tolerance factor and phase stability |
| I/Br | I/(Br + ) | Halide ratio: controls bandgap via valence band maximum; primary
– trade-off |
| Optical Depth | Pero_th2 | Quadratic thickness: Beer–Lambert absorption gain saturation |
| Vol. Recomb. | Pero_th × (MA + FA) | Thickness–composition interaction: volumetric recombination scaling |
Table 4.
Task-specific head configurations.
Table 4.
Task-specific head configurations.
| Target | Architecture | Hidden Layers | Physical Rationale |
|---|
| 128 → 64 → 1 | 2 (deepest) | Non-linear recombination physics with exponential dependence |
| 32 → 1 | 1 (shallowest) | Quasi-linear response to thickness and optical absorption |
| FF | 64 → 1 | 1 (moderate) | Non-linear resistance-related quantities |
| PCE | 32 → 1 | 1 (lightweight) | Multiplicative composite of the other three; minimal capacity needed |
Table 5.
Feature construction ablation (multi-task model, 5-fold CV).
Table 5.
Feature construction ablation (multi-task model, 5-fold CV).
| Configuration | | | FF | PCE | Cons. MAE |
|---|
| Raw Only (23 feat.) | | | | | 0.277 |
| Full (29 feat.) | | | | | 0.254 |
| No Interactions | | | | | 0.253 |
| No Ratios | | | | | 0.256 |
| No Optical Depth | | | | | 0.246 |
| No Vol. Recomb. | | | | | 0.238 |
Table 6.
Loss weighting sensitivity (5-fold CV).
Table 6.
Loss weighting sensitivity (5-fold CV).
| Config | | RMSE | | RMSE | FF | FF RMSE | PCE | Cons. MAE |
|---|
| 0.996 | 0.014 | 0.998 | 0.353 | 0.994 | 1.117 | 0.998 | 0.222 |
| * | 0.996 | 0.014 | 0.998 | 0.377 | 0.995 | 1.084 | 0.997 | 0.254 |
| 0.996 | 0.015 | 0.997 | 0.397 | 0.994 | 1.124 | 0.997 | 0.274 |
| 0.996 | 0.015 | 0.997 | 0.407 | 0.994 | 1.097 | 0.997 | 0.291 |
| 0.996 | 0.014 | 0.998 | 0.362 | 0.995 | 1.057 | 0.997 | 0.256 |
| 0.996 | 0.014 | 0.998 | 0.356 | 0.994 | 1.113 | 0.997 | 0.255 |
Table 7.
5-fold CV prediction accuracy (multi-task vs. single-task neural network).
Table 7.
5-fold CV prediction accuracy (multi-task vs. single-task neural network).
| Target | MT (Mean ± Std) | ST (Mean ± Std) | | Cohen’s d |
|---|
| | | | 3.58 |
| | | | 1.72 |
| FF | | | | 1.33 |
| PCE | | | | 8.93 |
Table 8.
Per-fold consistency MAE (PCE units).
Table 8.
Per-fold consistency MAE (PCE units).
| Model | Fold 1 | Fold 2 | Fold 3 | Fold 4 | Fold 5 | Mean |
|---|
| Multi-Task | 0.249 | 0.232 | 0.303 | 0.269 | 0.266 | 0.264 |
| Single-Task | 2.002 | 1.412 | 2.424 | 2.235 | 1.279 | 1.871 |
| Ratio (ST/MT) | 8.0× | 6.1× | 8.0× | 8.3× | 4.8× | 7.1× |
Table 9.
Physical consistency threshold analysis (multi-task vs. single-task).
Table 9.
Physical consistency threshold analysis (multi-task vs. single-task).
| Model | Cons. MAE | Std | Fraction > 2 PCE-Unit Threshold |
|---|
| Multi-Task | 0.264 | 0.343 | 0.014% (1 sample) |
| Single-Task | 1.871 | 2.487 | 36.5% (2616 samples) |
Table 10.
Full baseline comparison (5-fold CV).
Table 10.
Full baseline comparison (5-fold CV).
| Model | | | FF | PCE | Cons. MAE |
|---|
| Random Forest | | | | | |
| XGBoost | | | | | |
| CatBoost | | | | | |
| MH-NN (Ours) | | | | | |
Table 11.
Statistical significance testing (multi-task vs. single-task neural network).
Table 11.
Statistical significance testing (multi-task vs. single-task neural network).
| Target | MT | ST | Paired t p | Wilcoxon p | Cohen’s d |
|---|
| 0.996 | 0.961 | 0.001 | 0.063 | 3.58 |
| 0.998 | 0.963 | 0.018 | 0.063 | 1.72 |
| FF | 0.994 | 0.617 | 0.041 | 0.063 | 1.33 |
| PCE | 0.997 | 0.872 | <0.001 | 0.063 | 8.93 |
Table 12.
Per-composition (MT − ST) and sample sizes.
Table 12.
Per-composition (MT − ST) and sample sizes.
| Composition | N | | | FF | PCE | Note |
|---|
| CsPbI3 | 1008 | −0.000 | +0.000 | −0.000 | +0.001 | Largest comp. |
| MAPbI3 | 1001 | +0.001 | −0.001 | −0.002 | −0.001 | |
| FAPbI3 | 903 | +0.001 | +0.001 | +0.006 | +0.000 | |
| Cs0.15FA0.85 | 974 | +0.001 | −0.000 | +0.007 | +0.001 | |
| Cs0MA0.05FA0.95 | 420 | +0.081 | −0.004 | −0.004 | −0.006 | Largest gain |
| Cs0.1MA0.7FA0.2 | 490 | +0.009 | +0.002 | +0.005 | +0.013 | Lowest |
| Cs0MA0.3FA0.7 | 420 | +0.009 | −0.004 | +0.014 | −0.004 | Largest FF gain |
| Cs0MA0.15FA0.85 | 560 | +0.023 | +0.002 | +0.004 | +0.000 | |
| Cs0MA0.5FA0.5 | 350 | +0.017 | −0.001 | +0.002 | −0.001 | |
| Cs0MA0.7FA0.3 | 350 | +0.005 | −0.001 | +0.002 | +0.002 | |
| Cs0.05MA0.14FA0.81 | 350 | +0.009 | −0.009 | +0.001 | −0.003 | |
| Cs0.05MA0.16FA0.79 | 350 | −0.030 | −0.009 | +0.001 | +0.000 | Only regression |
Table 13.
Per-ETL (MT − ST) by electron transport layer.
Table 13.
Per-ETL (MT − ST) by electron transport layer.
| ETL | | | | FF | PCE | Mean |
|---|
| 1 | 77 | +0.002 | +0.002 | −0.020 | +0.005 | 0.007 |
| 2 | 226 | +0.001 | +0.001 | +0.005 | +0.000 | 0.002 |
| 3 | 212 | +0.001 | +0.000 | +0.002 | +0.000 | 0.001 |
| 4 | 223 | −0.001 | +0.000 | +0.002 | −0.000 | 0.001 |
| 5 | 133 | +0.001 | −0.000 | +0.001 | +0.006 | 0.002 |
| 6 | 219 | +0.000 | −0.001 | +0.004 | +0.000 | 0.001 |
| 7 | 215 | +0.002 | −0.001 | +0.003 | −0.001 | 0.002 |
| 8 | 131 | +0.002 | +0.000 | −0.001 | +0.001 | 0.001 |
Table 14.
Permutation importance (top 10 features, 50 repetitions).
Table 14.
Permutation importance (top 10 features, 50 repetitions).
| Feature | Imp. | Imp. | FF Imp. | PCE Imp. |
|---|
| I/Br (ratio) | 0.631 | 0.364 | 0.208 | 0.259 |
| FA | 0.371 | 0.557 | 0.338 | 0.523 |
| I | 0.501 | 0.405 | 0.194 | 0.227 |
| ETL7 | 0.114 | 0.519 | 0.323 | 0.485 |
| Cs | 0.287 | 0.214 | 0.289 | 0.160 |
| Cs/FA (ratio) | 0.235 | 0.251 | 0.167 | 0.236 |
| ETL2 | 0.041 | 0.222 | 0.161 | 0.315 |
| ETL6 | 0.044 | 0.209 | 0.178 | 0.302 |
| ETL3 | 0.039 | 0.205 | 0.143 | 0.283 |
| MA | 0.110 | 0.099 | 0.218 | 0.176 |
Table 15.
Hyperparameter sensitivity (5-fold CV ).
Table 15.
Hyperparameter sensitivity (5-fold CV ).
| Study | Configuration | | | FF | PCE |
|---|
| Width | | 0.993 | 0.994 | 0.987 | 0.993 |
| 0.996 | 0.997 | 0.993 | 0.996 |
| (Ours) | 0.996 | 0.998 | 0.994 | 0.998 |
| 0.996 | 0.998 | 0.995 | 0.998 |
| Depth | 1 layer | 0.996 | 0.998 | 0.995 | 0.998 |
| 2 layers (Ours) | 0.996 | 0.998 | 0.995 | 0.998 |
| 3 layers | 0.996 | 0.998 | 0.995 | 0.997 |
| 4 layers | 0.996 | 0.997 | 0.995 | 0.997 |
| Dropout | 0.0 | 0.992 | 0.994 | 0.989 | 0.993 |
| 0.1 | 0.996 | 0.998 | 0.995 | 0.998 |
| 0.2 (Ours) | 0.996 | 0.997 | 0.994 | 0.997 |
| 0.3 | 0.996 | 0.998 | 0.994 | 0.997 |
| 0.5 | 0.994 | 0.996 | 0.992 | 0.996 |
| Batch Size | 16 | 0.996 | 0.997 | 0.992 | 0.996 |
| 32 | 0.996 | 0.998 | 0.994 | 0.997 |
| 64 (Ours) | 0.996 | 0.998 | 0.994 | 0.997 |
| 128 | 0.996 | 0.997 | 0.995 | 0.997 |
| 256 | 0.995 | 0.997 | 0.994 | 0.996 |
Table 16.
Post-processing algebraic correction baseline. The corrected single-task model replaces the independently predicted PCE with . Consistency error is eliminated by construction, but PCE degrades due to compounded prediction errors from the three upstream single-task models.
Table 16.
Post-processing algebraic correction baseline. The corrected single-task model replaces the independently predicted PCE with . Consistency error is eliminated by construction, but PCE degrades due to compounded prediction errors from the three upstream single-task models.
| Fold | ST PCE | STcorr PCE | MT PCE | ST Cons. | STcorr Cons. | MT Cons. |
|---|
| 1 | 0.927 | 0.822 | 0.997 | 2.264 | 0.000 | 0.209 |
| 2 | 0.905 | 0.900 | 0.998 | 1.291 | 0.000 | 0.239 |
| 3 | 0.945 | 0.942 | 0.998 | 0.872 | 0.000 | 0.246 |
| 4 | 0.900 | 0.800 | 0.997 | 2.090 | 0.000 | 0.242 |
| 5 | 0.884 | 0.894 | 0.998 | 1.325 | 0.000 | 0.243 |
| Mean ± std | | | | | | |
Table 17.
Kendall learned-weight vs. manual configuration (5-fold CV). Learned weights are reported as effective precision () and normalized relative to .
Table 17.
Kendall learned-weight vs. manual configuration (5-fold CV). Learned weights are reported as effective precision () and normalized relative to .
| Target | Kendall | Manual | | Learned | Norm. |
|---|
| | | | | 1.25 |
| | | | | 1.00 |
| FF | | | | | 0.45 |
| PCE | | | | | 0.85 |
| Cons. MAE | | | | — | — |
Table 18.
Leave-one-composition-out (LOCO) results. Multi-task on held-out compositions not seen during training. Negative indicates predictions worse than the target mean. Single-task results are similarly catastrophic.
Table 18.
Leave-one-composition-out (LOCO) results. Multi-task on held-out compositions not seen during training. Negative indicates predictions worse than the target mean. Single-task results are similarly catastrophic.
| Held-Out Composition | | | | FF | PCE |
|---|
| CsPbI3 | 1008 | | | | |
| MAPbI3 | 1001 | | | | |
| FAPbI3 | 903 | | | | |
| Cs0.15FA0.85PbI3 | 974 | | | | |