MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors
Abstract
1. Introduction
- We formulate surface material assignment for 3D assets as a 2D material prediction and multi-view fusion problem, exploiting the relationship between visual semantics and material categories.
- We construct the MIO++ dataset, containing 115,542 single-object images from diverse camera angles, 12 object themes, and 30 fine-grained material categories with class-specific PBR assignments.
- We introduce MaterialSeg3D, which aggregates multi-view material predictions in UV space through weighted voting and region unification and converts the resulting labels to roughness and metallic maps.
- We analyze how irregular topology in AI-generated assets can affect PBR appearance and introduce an auxiliary plane-detection and simplification stage to improve planar surface quality before material assignment.
2. Related Works
2.1. 3D Asset Generation
2.2. Surface Material Generation
2.3. Existing 3D and 2D Datasets
3. Significance of Material for PBR
4. MIO++ Dataset
4.1. Consideration for Dataset Design
4.2. Data Collection and Annotation
4.3. Dataset Distribution
5. Method
5.1. Material Segmentation
5.2. MaterialSeg3D
5.3. Plane Simplification on AI-Generated Contents
6. Experiments
6.1. Implementations and Evaluations
6.2. Compared with Previous Work
6.3. Weighted Voting and Region Unification
6.4. Limitations and Open Issues
7. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A. Additional Visualization of Weighted Voting and Region Unification

References
- Hu, X.; Wang, Y.; Fan, L.; Fan, J.; Peng, J.; Lei, Z.; Li, Q.; Zhang, Z. Semantic anything in 3d gaussians. arXiv 2024, arXiv:2401.17857. [Google Scholar]
- Luo, S.; Hu, W. Diffusion probabilistic models for 3d point-cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 2837–2845. [Google Scholar]
- Achlioptas, P.; Diamanti, O.; Mitliagkas, I.; Guibas, L. Learning representations and generative models for 3d point clouds. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; pp. 40–49. [Google Scholar]
- Yang, G.; Huang, X.; Hao, Z.; Liu, M.Y.; Belongie, S.; Hariharan, B. Pointflow: 3d point-cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4541–4550. [Google Scholar]
- Zhou, L.; Du, Y.; Wu, J. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 5826–5835. [Google Scholar]
- Mescheder, L.; Oechsle, M.; Niemeyer, M.; Nowozin, S.; Geiger, A. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 4460–4470. [Google Scholar]
- Chen, Z.; Zhang, H. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 5939–5948. [Google Scholar]
- Zhang, S.H.; Guo, Y.C.; Gu, Q.W. Sketch2model: View-aware 3d modeling from single free-hand sketches. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 6012–6021. [Google Scholar]
- Qian, G.; Mai, J.; Hamdi, A.; Ren, J.; Siarohin, A.; Li, B.; Lee, H.Y.; Skorokhodov, I.; Wonka, P.; Tulyakov, S.; et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors. arXiv 2023, arXiv:2306.17843. [Google Scholar]
- Gao, J.; Shen, T.; Wang, Z.; Chen, W.; Yin, K.; Li, D.; Litany, O.; Gojcic, Z.; Fidler, S. Get3d: A generative model of high quality 3d textured shapes learned from images. Adv. Neural Inf. Process. Syst. 2022, 35, 31841–31854. [Google Scholar] [CrossRef] [Scilit]
- Liu, M.; Xu, C.; Jin, H.; Chen, L.; Xu, Z.; Su, H. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. arXiv 2023, arXiv:2306.16928. [Google Scholar]
- Long, X.; Guo, Y.C.; Lin, C.; Liu, Y.; Dou, Z.; Liu, L.; Ma, Y.; Zhang, S.H.; Habermann, M.; Theobalt, C.; et al. Wonder3d: Single image to 3d using cross-domain diffusion. arXiv 2023, arXiv:2310.15008. [Google Scholar]
- Tochilkin, D.; Pankratz, D.; Liu, Z.; Huang, Z.; Letts, A.; Li, Y.; Liang, D.; Laforte, C.; Jampani, V.; Cao, Y.P. Triposr: Fast 3d object reconstruction from a single image. arXiv 2024, arXiv:2403.02151. [Google Scholar]
- He, Z.; Wang, T. OpenLRM: Open-Source Large Reconstruction Models. Available online: https://github.com/3DTopia/OpenLRM (accessed on 26 May 2026).
- Upchurch, P.; Niu, R. A dense material segmentation dataset for indoor and outdoor scene parsing. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 450–466. [Google Scholar]
- Bell, S.; Upchurch, P.; Snavely, N.; Bala, K. Material recognition in the wild with the materials in context database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3479–3487. [Google Scholar]
- Li, Z.; Gan, R.; Luo, C.; Wang, Y.; Liu, J.; Zhu, Z.; Li, Q.; Yin, X.; Zhang, M.; Zhang, Z.; et al. Materialseg3d: Segmenting dense materials from 2d priors for 3d assets. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, 28 October–1 November 2024; pp. 370–379. [Google Scholar]
- Wu, J.; Zhang, C.; Xue, T.; Freeman, B.; Tenenbaum, J. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Adv. Neural Inf. Process. Syst. 2016, 29, 82–90. [Google Scholar]
- Smith, E.J.; Meger, D. Improved adversarial systems for 3d object generation and reconstruction. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2017; pp. 87–96. [Google Scholar]
- Xie, J.; Zheng, Z.; Gao, R.; Wang, W.; Zhu, S.C.; Wu, Y.N. Learning descriptor networks for 3d shape synthesis and analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8629–8638. [Google Scholar]
- Gadelha, M.; Maji, S.; Wang, R. 3d shape induction from 2d views of multiple objects. In Proceedings of the 2017 International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2017; pp. 402–411. [Google Scholar]
- Henzler, P.; Mitra, N.J.; Ritschel, T. Escaping plato’s cave: 3d shape from adversarial rendering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9984–9993. [Google Scholar]
- Lunz, S.; Li, Y.; Fitzgibbon, A.; Kushman, N. Inverse graphics gan: Learning to generate 3d shapes from unstructured 2d data. arXiv 2020, arXiv:2002.12674. [Google Scholar]
- Liu, Y.; Luo, C.; Fan, L.; Wang, N.; Peng, J.; Zhang, Z. CityGaussian: Real-Time High-Quality Large-Scale Scene Rendering with Gaussians. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 265–282. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Luo, C.; Mao, Z.; Peng, J.; Zhang, Z. CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025. [Google Scholar]
- Zhou, M.; Hou, J.; Luo, C.; Wang, Y.; Zhang, Z.; Peng, J. Scenex: Procedural controllable large-scale scene generation via large-language models. arXiv 2024, arXiv:2403.15698. [Google Scholar]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Commun. ACM 2022, 65, 99–106. [Google Scholar] [CrossRef] [Scilit]
- Jain, A.; Mildenhall, B.; Barron, J.T.; Abbeel, P.; Poole, B. Zero-shot text-guided object generation with dream fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 867–876. [Google Scholar]
- Mohammad Khalid, N.; Xie, T.; Belilovsky, E.; Popa, T. CLIP-Mesh: Generating Textured Meshes from Text Using Pretrained Image-Text Models. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers, Daegu, Republic of Korea, 6–9 December 2022; pp. 25:1–25:8. [Google Scholar] [CrossRef] [Scilit]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
- Poole, B.; Jain, A.; Barron, J.T.; Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. arXiv 2022, arXiv:2209.14988. [Google Scholar]
- Lin, C.H.; Gao, J.; Tang, L.; Takikawa, T.; Zeng, X.; Huang, X.; Kreis, K.; Fidler, S.; Liu, M.Y.; Lin, T.Y. Magic3d: High-resolution text-to-3d content creation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 300–309. [Google Scholar]
- Abdal, R.; Lee, H.Y.; Zhu, P.; Chai, M.; Siarohin, A.; Wonka, P.; Tulyakov, S. 3davatargan: Bridging domains for personalized editable avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 4552–4562. [Google Scholar]
- Chan, E.R.; Lin, C.Z.; Chan, M.A.; Nagano, K.; Pan, B.; De Mello, S.; Gallo, O.; Guibas, L.J.; Tremblay, J.; Khamis, S.; et al. Efficient geometry-aware 3D generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16123–16133. [Google Scholar]
- Chan, E.R.; Monteiro, M.; Kellnhofer, P.; Wu, J.; Wetzstein, G. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 5799–5809. [Google Scholar]
- Gu, J.; Liu, L.; Wang, P.; Theobalt, C. Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis. arXiv 2021, arXiv:2110.08985. [Google Scholar]
- Niemeyer, M.; Geiger, A. Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 11453–11464. [Google Scholar]
- Or-El, R.; Luo, X.; Shan, M.; Shechtman, E.; Park, J.J.; Kemelmacher-Shlizerman, I. Stylesdf: High-resolution 3d-consistent image and geometry generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 13503–13513. [Google Scholar]
- Skorokhodov, I.; Siarohin, A.; Xu, Y.; Ren, J.; Lee, H.Y.; Wonka, P.; Tulyakov, S. 3d generation on imagenet. arXiv 2023, arXiv:2303.01416. [Google Scholar]
- Xu, Y.; Chai, M.; Shi, Z.; Peng, S.; Skorokhodov, I.; Siarohin, A.; Yang, C.; Shen, Y.; Lee, H.Y.; Zhou, B.; et al. DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 4402–4412. [Google Scholar]
- Asselin, L.P.; Laurendeau, D.; Lalonde, J.F. Deep SVBRDF estimation on real materials. In Proceedings of the 2020 International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2020; pp. 1157–1166. [Google Scholar]
- Deschaintre, V.; Lin, Y.; Ghosh, A. Deep polarization imaging for 3D shape and SVBRDF acquisition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 15567–15576. [Google Scholar]
- Deschaintre, V.; Aittala, M.; Durand, F.; Drettakis, G.; Bousseau, A. Single-image svbrdf capture with a rendering-aware deep network. ACM Trans. Graph. 2018, 37, 128. [Google Scholar] [CrossRef] [Scilit]
- Martin, R.; Roullier, A.; Rouffet, R.; Kaiser, A.; Boubekeur, T. MaterIA: Single Image High-Resolution Material Capture in the Wild. Comput. Graph. Forum 2022, 41, 163–177. [Google Scholar] [CrossRef] [Scilit]
- Gao, D.; Li, X.; Dong, Y.; Peers, P.; Xu, K.; Tong, X. Deep Inverse Rendering for High-Resolution SVBRDF Estimation from an Arbitrary Number of Images. ACM Trans. Graph. 2019, 38, 134:1–134:15. [Google Scholar] [CrossRef] [Scilit]
- Deschaintre, V.; Drettakis, G.; Bousseau, A. Guided Fine-Tuning for Large-Scale Material Transfer. Comput. Graph. Forum 2020, 39, 91–105. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Dong, Y.; Peers, P.; Tong, X. Modeling surface appearance from a single photograph using self-augmented convolutional neural networks. ACM Trans. Graph. 2017, 36, 45. [Google Scholar] [CrossRef] [Scilit]
- Vecchio, G.; Palazzo, S.; Spampinato, C. SurfaceNet: Adversarial SVBRDF Estimation from a Single Image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 12840–12848. [Google Scholar]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. arXiv 2023, arXiv:2304.02643. [Google Scholar]
- Vecchio, G.; Sortino, R.; Palazzo, S.; Spampinato, C. Matfuse: Controllable material generation with diffusion models. arXiv 2023, arXiv:2308.11408. [Google Scholar]
- Sartor, S.; Peers, P. Matfusion: A generative diffusion model for svbrdf capture. In Proceedings of the SIGGRAPH Asia 2023 Conference Papers, Sydney, Australia, 12–15 December 2023; pp. 86:1–86:10. [Google Scholar]
- Vecchio, G.; Martin, R.; Roullier, A.; Kaiser, A.; Rouffet, R.; Deschaintre, V.; Boubekeur, T. ControlMat: A Controlled Generative Approach to Material Capture. arXiv 2023, arXiv:2309.01700. [Google Scholar]
- Chen, R.; Chen, Y.; Jiao, N.; Jia, K. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. arXiv 2023, arXiv:2303.13873. [Google Scholar]
- Yeh, Y.Y.; Li, Z.; Hold-Geoffroy, Y.; Zhu, R.; Xu, Z.; Hašan, M.; Sunkavalli, K.; Chandraker, M. Photoscene: Photorealistic material and lighting transfer for indoor scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18562–18571. [Google Scholar]
- Yuan, L.; Yan, D.; Saito, S.; Fujishiro, I. DiffMat: Latent diffusion models for image-guided material generation. Vis. Inform. 2024, 8, 6–14. [Google Scholar] [CrossRef] [Scilit]
- Lopes, I.; Pizzati, F.; de Charette, R. Material Palette: Extraction of Materials from a Single Image. arXiv 2023, arXiv:2311.17060. [Google Scholar]
- Ceylan, D.; Deschaintre, V.; Groueix, T.; Martin, R.; Huang, C.H.; Rouffet, R.; Kim, V.; Lassagne, G. MatAtlas: Text-driven Consistent Geometry Texturing and Material Assignment. arXiv 2024, arXiv:2404.02899. [Google Scholar]
- Zhang, Y.; Tu, Z.; Lian, W.; Hu, Y.; Sun, S.; Xiao, Y.; Cheng, Y. BANet: Bidirectional Feature Aggregation and Adaptive Multi-Scene Perception-Based Lane Detection for Autonomous Driving. IEEE Trans. Intell. Transp. Syst. 2026, 27, 9713–9723. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Peng, J.; Zhang, G.; Luo, C.; Xu, S.; Zhang, M.; Zhang, Z. FurniScene: A Large-Scale 3D Room Dataset with Intricate Furnishing Scenes. Int. J. Comput. Vis. 2026, 134, 125. [Google Scholar] [CrossRef] [Scilit]
- Deitke, M.; Schwenk, D.; Salvador, J.; Weihs, L.; Michel, O.; VanderBilt, E.; Schmidt, L.; Ehsani, K.; Kembhavi, A.; Farhadi, A. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 13142–13153. [Google Scholar]
- Deitke, M.; Liu, R.; Wallingford, M.; Ngo, H.; Michel, O.; Kusupati, A.; Fan, A.; Laforte, C.; Voleti, V.; Gadre, S.Y.; et al. Objaverse-xl: A universe of 10m+ 3d objects. arXiv 2023, arXiv:2307.05663. [Google Scholar]
- Kasper, A.; Xue, Z.; Dillmann, R. The kit object models database: An object model database for object recognition, localization and manipulation in service robotics. Int. J. Robot. Res. 2012, 31, 927–934. [Google Scholar] [CrossRef] [Scilit]
- Calli, B.; Walsman, A.; Singh, A.; Srinivasa, S.; Abbeel, P.; Dollar, A.M. Benchmarking in manipulation research: The ycb object and model set and benchmarking protocols. arXiv 2015, arXiv:1502.03143. [Google Scholar]
- Singh, A.; Sha, J.; Narayan, K.S.; Achim, T.; Abbeel, P. Bigbird: A large-scale 3d database of object instances. In Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2014; pp. 509–516. [Google Scholar]
- Sun, X.; Wu, J.; Zhang, X.; Zhang, Z.; Zhang, C.; Xue, T.; Tenenbaum, J.B.; Freeman, W.T. Pix3d: Dataset and methods for single-image 3d shape modeling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 2974–2983. [Google Scholar]
- Downs, L.; Francis, A.; Koenig, N.; Kinman, B.; Hickman, R.; Reymann, K.; McHugh, T.B.; Vanhoucke, V. Google scanned objects: A high-quality dataset of 3d scanned household items. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2022; pp. 2553–2560. [Google Scholar]
- Park, K.; Rematas, K.; Farhadi, A.; Seitz, S.M. Photoshape: Photorealistic materials for large-scale shape collections. arXiv 2018, arXiv:1809.09761. [Google Scholar]
- Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; Xiao, J. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 1912–1920. [Google Scholar]
- Wu, R.; Xiao, C.; Zheng, C. Deepcad: A deep generative network for computer-aided design models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 6772–6782. [Google Scholar]
- Koch, S.; Matveev, A.; Jiang, Z.; Williams, F.; Artemov, A.; Burnaev, E.; Alexa, M.; Zorin, D.; Panozzo, D. ABC: A Big CAD Model Dataset for Geometric Deep Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 9601–9611. [Google Scholar]
- Bell, S.; Upchurch, P.; Snavely, N.; Bala, K. OpenSurfaces: A richly annotated catalog of surface appearance. ACM Trans. Graph. 2013, 32, 111. [Google Scholar]
- images.cv. CV Image Dataset. 2024. Available online: https://images.cv (accessed on 26 May 2026).
- Yang, L.; Luo, P.; Loy, C.C.; Tang, X. A large-scale car dataset for fine-grained categorization and verification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3973–3981. [Google Scholar]
- Aubry, M.; Maturana, D.; Efros, A.A.; Russell, B.C.; Sivic, J. Seeing 3d chairs: Exemplar part-based 2d-3d alignment using a large dataset of cad models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 3762–3769. [Google Scholar]
- Wu, T.; Li, Z.; Yang, S.; Zhang, P.; Pan, X.; Wang, J.; Lin, D.; Liu, Z. HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image. In Proceedings of the SIGGRAPH Asia 2023 Conference Papers, Sydney, Australia, 12–15 December 2023; pp. 53:1–53:10. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Xiao, T.; Liu, Y.; Zhou, B.; Jiang, Y.; Sun, J. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 418–434. [Google Scholar]
- Contributors, M. MMEngine: OpenMMLab Foundational Library for Training Deep Learning Models. Available online: https://github.com/open-mmlab/mmengine (accessed on 26 May 2026).
- Chen, D.Z.; Siddiqui, Y.; Lee, H.Y.; Tulyakov, S.; Nießner, M. Text2tex: Text-driven texture synthesis via diffusion models. arXiv 2023, arXiv:2303.11396. [Google Scholar]
- Gan, R.; Peng, J.; Liu, Y.; Luo, C.; Li, Q.; Zhang, Z. GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation. arXiv 2025, arXiv:2510.17095. [Google Scholar]
- Yin, W.; Zhang, C.; Chen, H.; Cai, Z.; Yu, G.; Wang, K.; Chen, X.; Shen, C. Metric3d: Towards zero-shot metric 3d prediction from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 9043–9053. [Google Scholar]
- Derpanis, K.G. Overview of the RANSAC Algorithm. Image 2010, 4, 2–3. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. arXiv 2018, arXiv:1711.05101. [Google Scholar]
- Contributors, M. MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark. 2020. Available online: https://github.com/open-mmlab/mmsegmentation (accessed on 1 June 2026).
- Liu, Z.; Mao, H.; Wu, C.Y.; Feichtenhofer, C.; Darrell, T.; Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 11976–11986. [Google Scholar]
- Sun, K.; Zhao, Y.; Jiang, B.; Cheng, T.; Xiao, B.; Liu, D.; Mu, Y.; Wang, X.; Liu, W.; Wang, J. High-resolution representations for labeling pixels and regions. arXiv 2019, arXiv:1904.04514. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16000–16009. [Google Scholar]














| Aspect | ACM MM 2024 Version | This Journal Version |
|---|---|---|
| Dataset | MIO: 23,062 images | MIO++: 115,542 images |
| Object themes | 5 | 12 |
| Material-label space | 14 categories | 30 fine-grained categories |
| PBR assignment | 14 class-level settings | 30 independent class-level settings |
| AI-generated meshes | Original material workflow | Added plane-oriented preprocessing |
| Evaluation | Original MIO evaluation | MIO++ analysis and expanded 3D evaluations |
| Material Label | Number | Material Label | Number |
|---|---|---|---|
| metal | 935 | brick | 186 |
| wood | 842 | porcelain | 163 |
| plastic | 768 | clay terracotta | 154 |
| glass | 712 | concrete | 152 |
| paint | 626 | nylon | 75 |
| rubber | 524 | rusty metal | 53 |
| leather | 437 | stone | 46 |
| fabric | 391 | bone | 25 |
| fruit&leaf | 273 | bamboo | 22 |
| flower | 252 | others | 181 |
| Label Id. | Material Label | MIO | MIO++ |
|---|---|---|---|
| 1 | rubber | 5324 | 7244 |
| 2 | smooth metal | 8946 | 18,047 |
| 3 | matte metal | ↑ | 15,429 |
| 4 | rusty metal | ↑ | 1560 |
| 5 | paint | 3496 | 4003 |
| 6 | glass | 5802 | 19,637 |
| 7 | leather | 3417 | 3496 |
| 8 | smooth wood | 7088 | 11,691 |
| 9 | rough wood | ↑ | 6156 |
| 10 | rough fabric | 3373 | 12,270 |
| 11 | smooth fabric | ↑ | 1283 |
| 12 | plush fabric | ↑ | 2199 |
| 13 | smooth plastic | 6928 | 11,807 |
| 14 | rough plastic | ↑ | 10,939 |
| 15 | brick | 1017 | 2689 |
| 16 | concrete | 794 | 4306 |
| 17 | clay | 910 | 2290 |
| 18 | ceramic | 921 | 6604 |
| 19 | flower | 1677 | 2577 |
| 20 | fruit | 1742 | 4758 |
| 21 | leaf | ↑ | 3020 |
| 22 | meat | - | 3345 |
| 23 | gel | - | 2438 |
| 24 | cream | - | 1957 |
| 25 | biscuit | - | 5287 |
| 26 | shell | - | 756 |
| 27 | water | - | 1596 |
| 28 | jade | - | 519 |
| 29 | pearl | - | 212 |
| 30 | plaster | - | 2014 |
| Classname (Abbr.) | Rendered Image | Real Image | Total |
|---|---|---|---|
| Furniture (fur.) | 4152 | 5455 | 9607 |
| Cars (car) | 1935 | 4117 | 6052 |
| Buildings (bui.) | 418 | 1752 | 2170 |
| Musical Instrument (ins.) | 627 | 1637 | 2264 |
| Plants (pla.) | 552 | 2417 | 2969 |
| Classname (Abbr.) | MIO | New in MIO++ | Total |
|---|---|---|---|
| Furniture (fur.) | 9607 | 2499 | 12,106 |
| Car (car) | 6052 | 5558 | 11,610 |
| Building (bui.) | 2170 | 6488 | 8658 |
| Musical Instrument (ins.) | 2264 | 8337 | 10,601 |
| Statue (sta.) | - | 2400 | 2400 |
| Household appliance (hou.) | - | 11,991 | 11,991 |
| Kitchen ware (kit.) | - | 11,992 | 11,992 |
| Weapon and tool (wea.) | - | 11,891 | 11,891 |
| Clothing (clo.) | - | 13,906 | 13,906 |
| Food (foo.) | - | 14,970 | 14,970 |
| Jewelry (jew.) | - | 2808 | 2808 |
| Plants (pla.) | 2969 | - | 2969 |
| Total | 23,062 | 92,840 | 115,542 |
| bui. | car | clo. | hou. | foo. | jew. | kit. | mus. | sta. | fur. | wea. | pla. | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rubber | – | 73.57 | 73.84 | – | – | – | – | 8.46 | – | – | 59.98 | – |
| smooth metal | – | 55.78 | 41.81 | 41.67 | 11.06 | 70.78 | 52.73 | 76.67 | 23.41 | 19.91 | 42.02 | – |
| matte metal | 49.78 | – | 42.85 | 45.11 | – | 21.54 | 45.4 | 43.57 | 52.45 | 36.61 | 84.63 | – |
| rusty metal | 49.91 | – | – | 19.73 | – | – | 1.17 | – | 31.33 | – | 75.35 | – |
| paint | – | 79.93 | – | – | – | – | – | – | – | – | – | – |
| glass | 76.3 | 84.97 | 57.82 | 80.06 | 18.63 | 69.28 | 59.89 | – | – | 81.82 | – | – |
| leather | – | 52.25 | 66.27 | 45.24 | – | – | – | 61.63 | – | 62.26 | 26.74 | – |
| smooth wood | – | – | – | 51.35 | – | – | 73.86 | 74.27 | 43.55 | 66.08 | 48.76 | – |
| rough wood | 71.24 | – | 57.3 | – | – | – | 55.94 | 48.96 | 22.49 | 61.13 | 54.18 | 78.73 |
| rough fabric | 64.22 | – | 88.97 | 44.29 | – | 43.66 | 24.36 | – | – | 84.18 | 67.95 | – |
| smooth fabric | – | – | 29.48 | – | – | – | – | – | – | – | – | – |
| plush fabric | – | – | 47.57 | – | – | 22.34 | – | – | – | – | – | – |
| smooth plastic | – | 61.07 | 29.31 | 39.33 | – | – | 23.24 | 66.43 | – | 43.06 | 51.84 | – |
| rough plastic | – | 36.56 | 42.32 | 71.2 | – | – | 63.81 | 42.45 | – | 52.76 | 24.39 | – |
| brick | 76.73 | – | – | – | – | – | – | – | – | – | – | – |
| concrete | 80.27 | – | – | – | – | – | – | – | – | – | – | – |
| clay | 64.59 | – | – | – | – | – | 32.57 | – | 38.81 | – | – | 80.18 |
| ceramic | – | – | – | 17.84 | 45.72 | – | 77.25 | – | – | 47.35 | – | 86.48 |
| flower | – | – | – | – | – | – | 27.41 | – | – | – | – | 94.19 |
| fruit | – | – | – | – | 87.26 | – | 69.08 | – | – | – | – | – |
| leaf | 76.12 | – | – | – | 39.04 | – | 57.29 | 55.63 | – | 36.86 | – | 92.30 |
| meat | – | – | – | – | 67.25 | – | – | – | – | – | – | – |
| gel | – | – | – | – | 53.82 | – | – | – | – | – | – | – |
| cream | – | – | – | – | 72.54 | – | – | – | – | – | – | – |
| biscuit | – | – | – | 55.3 | 77.04 | – | 44.71 | – | – | – | – | – |
| shell | – | – | – | – | 75.57 | – | – | – | – | – | – | – |
| water | – | – | – | – | 43.21 | – | 37.18 | – | – | – | – | – |
| jade | – | – | – | – | – | 61.16 | – | – | – | – | – | – |
| pearl | – | – | – | – | – | 54.77 | – | – | – | – | – | – |
| plaster | – | – | – | – | – | – | – | – | 89.56 | – | – | – |
| mIoU | 67.68 | 62.59 | 52.50 | 46.47 | 53.74 | 49.08 | 46.62 | 53.12 | 43.09 | 53.82 | 53.584 | 86.38 |
| Method | MIO Dataset (%) | |||||
|---|---|---|---|---|---|---|
| car | fur. | bui. | ins. | pla. | mIoU | |
| ConvNeXt [85] | 71.03 | 74.85 | 69.33 | 72.40 | 76.72 | 72.87 |
| HRNet [86] | 75.71 | 79.94 | 76.37 | 80.14 | 81.35 | 78.70 |
| ViT [76] | 73.96 | 77.67 | 75.53 | 79.45 | 78.66 | 77.05 |
| Swin-T [87] | 75.09 | 79.04 | 78.45 | 80.92 | 81.40 | 78.98 |
| MAE [88] | 76.42 | 82.06 | 77.59 | 82.74 | 85.92 | 80.95 |
| Ours(MIO) | 81.83 | 85.22 | 81.76 | 84.39 | 86.38 | 83.92 |
| Ours(MIO++) | 85.14 | 87.35 | 84.49 | 88.82 | 87.97 | 86.75 |
| Method | Input Setting | CLIP Similarity↑ | PSNR↑ | SSIM↑ | |||
|---|---|---|---|---|---|---|---|
| Reference | Novel | Reference | Novel | Reference | Novel | ||
| Wonder3D [12] | Reconstructed | 0.85 | 0.84 | 16.06 | 15.83 | 0.78 | 0.75 |
| TripoSR [13] | 0.93 | 0.90 | 16.93 | 16.14 | 0.79 | 0.76 | |
| OpenLRM [14] | 0.92 | 0.87 | 16.30 | 15.37 | 0.77 | 0.76 | |
| Baseline | Predefined mesh + albedo | 0.93 | 0.93 | 16.28 | 16.30 | 0.79 | 0.78 |
| Baseline + Ours | 0.98 | 0.97 | 20.72 | 18.39 | 0.85 | 0.84 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Peng, J.; Gan, R.; Shen, S.; Li, Z.; Liu, Y.; Zhu, Z. MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors. Electronics 2026, 15, 3885. https://doi.org/10.3390/electronics15173885
Peng J, Gan R, Shen S, Li Z, Liu Y, Zhu Z. MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors. Electronics. 2026; 15(17):3885. https://doi.org/10.3390/electronics15173885
Chicago/Turabian StylePeng, Junran, Ruitong Gan, Silei Shen, Zongxing Li, Yan Liu, and Ziwei Zhu. 2026. "MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors" Electronics 15, no. 17: 3885. https://doi.org/10.3390/electronics15173885
APA StylePeng, J., Gan, R., Shen, S., Li, Z., Liu, Y., & Zhu, Z. (2026). MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors. Electronics, 15(17), 3885. https://doi.org/10.3390/electronics15173885

