Next Article in Journal
Analytical Approach to Uncertainty Evaluation of Exposure Index—Metric for Assessing Human Population Exposure Induced by a Wireless Cellular Network
Previous Article in Journal
Physics-Informed Transfer Learning for Cross-Condition State of Health Prediction of Lithium-Ion Batteries
Previous Article in Special Issue
HardMo++: A Large-Scale Hard-Case Dataset for Motion Capture
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors

1
School of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China
2
Shunde Innovation School, University of Science and Technology Beijing, Foshan 528399, China
3
Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China
4
School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China
5
Key Laboratory of Knowledge Automation for Industrial Processes of Ministry of Education, University of Science and Technology Beijing, Beijing 100083, China
6
Dawa Future (Beijing) Imaging Technology Co., Ltd., Beijing 100144, China
*
Authors to whom correspondence should be addressed.
These authors contributed equally to this work.
Electronics 2026, 15(17), 3885; https://doi.org/10.3390/electronics15173885 (registering DOI)
Submission received: 30 June 2026 / Revised: 15 August 2026 / Accepted: 21 August 2026 / Published: 28 August 2026

Abstract

Recent image diffusion models have enabled automatic 3D object creation from text or image guidance, but their 2D generative priors often bake illumination and shadow into textures, making relighting and physically based rendering (PBR) difficult. To address this issue, we propose MaterialSeg3D, a framework that predicts surface materials for 3D assets by leveraging 2D material semantics. Given a mesh and its albedo UV map, MaterialSeg3D renders multi-view images, performs material segmentation using a 2D prior model, projects the predictions back to UV space, and fuses them through weighted voting and region unification to obtain coherent material maps. To train the prior model, we construct MIO++, a large-scale single-object material segmentation dataset containing 115,542 images, 12 object themes, and 30 fine-grained material categories, substantially extending the previous MIO dataset. Each MIO++ material category is associated with an independent pair of roughness and metallic values for PBR assignment. We further observe that poor topology in AI-generated assets can degrade PBR quality even when plausible materials are assigned, and introduce a plane-simplification strategy as an auxiliary preprocessing step for such meshes. Experiments show that MIO++ improves material segmentation and that the resulting material maps support more consistent relightable renderings for both human-crafted and AI-generated 3D assets.

1. Introduction

3D asset creation is an important topic in computer graphics, with applications in virtual reality, augmented reality, games, and film production. A central objective is to construct assets that are visually plausible, artistically controllable, and compatible with different environments. Physically based rendering (PBR) is particularly important for this purpose because it models how light interacts with surface properties such as roughness and metallicity. Creating assets that satisfy PBR requirements remains labor-intensive in traditional production pipelines. Artists must not only construct suitable geometry and textures, but also define material properties that behave consistently under changing illumination. With the recent development of diffusion models, many methods have been proposed for efficient 3D asset generation, including approaches based on point clouds [1,2,3,4,5], implicit functions [6,7], and meshes [8,9,10,11,12,13,14]. These methods can generate 3D objects from textual or visual guidance, and many of them rely on powerful 2D generative priors. However, illumination and shadows can be baked into the generated appearance, so an RGB texture alone does not necessarily provide a disentangled material representation. Without explicit PBR material information, relighting in novel scenes remains difficult. Directly learning PBR material generation is also constrained by data availability. Many existing 3D assets provide geometry and appearance textures but lack complete roughness and metallic maps, while material information represented in UV space is expensive to annotate densely. This motivates us to transfer material reasoning to the better-supported 2D perception space and then project the predictions back to the 3D asset.
To address this problem, we build a large-scale material dataset covering common material types in 3D asset creation and propose a 3D material prediction workflow based on 2D priors. We follow the Disney principled bidirectional reflectance distribution function (BRDF) and use roughness and metalness as the primary physical material properties. These properties modulate the BRDF terms and enable a surface to respond differently under changing illumination. To design a dataset that is useful for PBR while remaining practical to annotate, we draw on the observation that material semantics can often be inferred from object appearance. Given reference images of an object, a modeler may infer, for example, that silver chair legs are metallic while a dark padded seat is likely leather or fabric. This motivates the use of 2D material segmentation as a source of material prior knowledge for 3D assets. Existing material-related segmentation datasets such as DMS [15] or MINC [16] only provide material labels for open scenes including multiple instances, which are less reliable in dealing with single-object component segmentation. More importantly, images in these datasets are taken from very limited camera poses, which makes it hard for models to learn the material prediction from all angles. With the motivation of establishing a database to construct 2D material prior knowledge for individual objects, we collect Materialized Individual Objects (MIO), a novel 2D single-object segmentation dataset consisting of dense material-semantic annotations of objects with intricate semantic classes and captured camera angles. Images are (a) collected from both real-world captures and 3D asset renderings, augmenting the prior knowledge from reality and easing the domain gap; (b) sampled with various camera angles including but not limited to top and side views; and (c) annotated and supervised by professional annotators. For each material class, we associate representative roughness and metalness values selected with guidance from experienced modelers. The original MIO dataset contains 23,062 images, five object themes, and 14 material categories. We substantially extend it to MIO++, which contains 115,542 images, 12 object themes, and 30 fine-grained material categories. In MIO++, materials with different PBR behavior are separated into finer categories, such as smooth metal, matte metal, and rusty metal. Importantly, the 30 categories are retained in the actual MaterialSeg3D pipeline and each category has its own roughness and metallic assignment.
Using MIO++, we develop MaterialSeg3D, a workflow for assigning surface materials to 3D objects. Given a geometry mesh and an albedo UV map, the method renders the asset from multiple camera poses, predicts material labels in each view, projects the predictions back to temporary UV maps, and combines them through weighted voting and region unification. The resulting material-label UV map is converted to roughness and metallic maps using the class-specific assignments defined for MIO++. As shown in Figure 1, this process transfers 2D material priors to a 3D asset and provides explicit PBR material maps that can be used for relighting.
We also observe that poor topology in AI-generated assets can reduce PBR quality even when a plausible material has been assigned. Small geometric irregularities on surfaces that should be planar may produce unstable normals and visually inconsistent reflections. For AI-generated assets, we therefore include an auxiliary plane-detection and simplification stage before applying the material workflow. This stage is intended as practical geometry preprocessing rather than as a general-purpose mesh reconstruction method.
This article is a revised and substantially expanded version of the conference paper entitled “MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets”, which was presented at the 32nd ACM International Conference on Multimedia (MM ’24), Melbourne, VIC, Australia, 28 October–1 November 2024 [17]. Compared with the conference version, this article introduces MaterialSeg3D++, a substantially expanded study with several major extensions. First, we expand the original MIO dataset from 23,062 images, five object themes, and 14 material classes to MIO++ with 115,542 images, 12 object themes, and 30 fine-grained material categories, yielding about a five-fold increase in dataset scale and a more than two-fold expansion in label space. Second, MIO++ introduces more fine-grained material categories based on roughness and metalness, and training on MIO++ improves material prediction performance, increasing mIoU from 83.92% to 86.75% after mapping fine-grained subclasses to the original meta-classes. Third, we extend the original workflow to AI-generated assets by adding a plane-simplification strategy to improve physically based rendering on poor-quality generated meshes. Finally, we provide additional quantitative and qualitative evaluations, including detailed validation on MIO++, comparisons with image-to-3D and texture generation methods, and more visual analyses under different lighting conditions.
For clarity, Table 1 summarizes these extensions side by side.
To summarize, the contributions of this paper are as follows:
  • We formulate surface material assignment for 3D assets as a 2D material prediction and multi-view fusion problem, exploiting the relationship between visual semantics and material categories.
  • We construct the MIO++ dataset, containing 115,542 single-object images from diverse camera angles, 12 object themes, and 30 fine-grained material categories with class-specific PBR assignments.
  • We introduce MaterialSeg3D, which aggregates multi-view material predictions in UV space through weighted voting and region unification and converts the resulting labels to roughness and metallic maps.
  • We analyze how irregular topology in AI-generated assets can affect PBR appearance and introduce an auxiliary plane-detection and simplification stage to improve planar surface quality before material assignment.

2. Related Works

2.1. 3D Asset Generation

Early methods in 3D asset generation adapted convolutional neural networks (CNNs) and generative adversarial networks (GANs) to generate 3D voxel grids [18,19,20,21,22,23]. Although straightforward, voxel representations can incur substantial memory and computational cost at high resolutions.
Subsequent research explored point-cloud representations [1,2,3,4,5] and implicit functions [6,7]. Gaussian-based reconstruction methods have further improved large-scale rendering and geometric accuracy [24,25], while controllable procedural frameworks have explored large-scale 3D scene generation [26]. Mesh-based generative models [8,9,10,11,12,13,14] are particularly convenient for standard graphics pipelines because their outputs can be directly rendered and edited in conventional engines.
Recent 3D generation methods also use text or image guidance together with Neural Radiance Fields (NeRFs) [27], CLIP-based alignment [28,29,30], or diffusion priors. DreamFusion [31] introduced score distillation sampling to optimize 3D representations from a diffusion model, while Magic3D [32] adopted a coarse-to-fine strategy. Other methods [33,34,35,36,37,38,39,40] combine learned 3D representations with powerful 2D generative priors. These advances improve geometry and appearance generation, but explicit PBR material assignment remains a distinct problem.

2.2. Surface Material Generation

PBR properties such as metalness and roughness determine how a 3D surface interacts with incident light and are therefore important for realistic relighting and reflection behavior.
Traditional material-acquisition methods often estimate physically based materials under controlled lighting and may require multi-view [41] or polarization [42] equipment. Learning-based approaches use synthetic data to train single-view Spatially Varying Bidirectional Reflectance Distribution Function (SVBRDF) prediction networks [43], sometimes combined with real single-view data [44,45] or specialized training strategies [46,47,48]. These methods address material capture directly, whereas our setting assumes an existing 3D asset and focuses on semantic material assignment over its surface.
Recent work has also investigated material segmentation and controllable generation of SVBRDF maps [49,50,51,52]. Fantasia3D [53] decouples geometry and appearance modeling and uses a BRDF-based representation for high-quality 3D content. PhotoScene [54] uses procedural material graphs, while DiffMat [55] and Material Palette [56] address material generation or extraction in related settings. MatAtlas [57] generates relightable textures and material assignments for 3D models from text guidance. Our work differs in that it predicts discrete material semantics from rendered views and fuses these predictions in the UV domain. More broadly, recent dense-perception methods have explored bidirectional feature aggregation and scene-adaptive perception to improve robustness across visual conditions [58]; these ideas are complementary to our view-level fusion setting, although they address a different task domain.

2.3. Existing 3D and 2D Datasets

Learning material priors requires sufficiently diverse training data. Several large-scale 3D datasets have recently been released [59]; representative examples include Objaverse-1.0 [60] and Objaverse-XL [61], with approximately 800,000 and more than 10 million assets, respectively. However, many Objaverse assets do not contain complete material information suitable for supervised PBR learning. Other datasets such as KIT [62], YCB [63], BigBIRD [64], and Pix3D [65] provide calibrated object models but are much smaller. Photorealistic object collections [66,67] and CAD datasets [68,69,70] also do not consistently provide the albedo and PBR material maps required by our setting.
Because 2D images are easier to collect, material recognition has access to larger 2D datasets. DMS [15] contains 44,560 indoor and outdoor images with annotations for 3.2 million dense segments. OpenSurfaces [71] provides annotations for 37 material categories on 19,000 residential images, and MINC [16] contains millions of annotated material samples. These datasets are valuable for general material recognition, but they are dominated by multi-object scene imagery and conventional viewpoints. Our application instead requires dense predictions on a single object rendered from many surrounding and top/oblique viewpoints.

3. Significance of Material for PBR

Creating high-quality materials in computer graphics is a demanding task that requires substantial expertise. Correct PBR material properties allow the same 3D asset to respond plausibly under different lighting conditions, whereas missing or inappropriate material parameters can produce unrealistic reflection and relighting behavior. Figure 2 illustrates this difference under a common lighting setup.
To assess whether existing 3D collections could directly support material learning, we examined more than 270,000 assets from Objaverse [61]. Only about 3k of the examined assets contained material information that was sufficiently complete for our initial experiments. This limited subset was not adequate for learning the material-semantic distribution needed by our pipeline.
Recent 3D generation and material generation methods have made progress toward relightable appearance [50,51,53,57]. In our experiments with generated assets, we nevertheless observe two practical issues that motivate the following analysis.
One issue is that a uniform or weakly differentiated PBR setting can ignore semantic differences between object parts. As shown in Figure 3a, the hammer handle and head receive the same PBR values even though they correspond to visibly different materials. A second issue is semantic inconsistency within a continuous object region. In Figure 3b, parts of the chair back are assigned metal even though the region is visually more consistent with a continuous material such as fabric or nylon.
These observations motivate the use of semantic material priors. As an exploratory check of this motivation, we surveyed 100 participants and showed each participant 10 images drawn from indoor and outdoor scenes, giving 1000 image–participant observations in total. Participants were asked to record the material types that they considered likely to occur in each image, and the reported frequencies were aggregated as shown in Table 2. The survey was exploratory rather than demographically stratified, and detailed recruitment or demographic metadata were not retained. It is used only to support the general observation that people can associate visible object regions with material categories, not to determine the numerical roughness or metallic values assigned to individual material classes; those PBR assignments were established separately with experienced 3D asset modelers, as described in Section 4.2.

4. MIO++ Dataset

4.1. Consideration for Dataset Design

In an early attempt to solve the material assignment problem, we explored material priors from existing 3D asset datasets such as Objaverse-1.0 and Objaverse-XL. We collected objects with available material UV maps and applied k-means clustering with the elbow method to group materials according to their roughness and metalness values. However, only a small fraction of the examined assets contained sufficiently complete and usable PBR information (about 3k out of 200k examined assets), and the resulting data were not adequate for training a robust material recognition model. We therefore constructed a dedicated material dataset instead of relying on these existing 3D assets.
Our pilot study suggests that 2D appearance provides useful prior information for material categorization. Rather than annotating material labels directly in UV space, we therefore formulate material prediction as pixel-wise segmentation in rendered or captured 2D views and later project those predictions back to the UV map. Different roughness and metallic properties can lead to substantially different appearances under the same illumination conditions, as illustrated in Figure 4.
To determine the material vocabulary, we first examined more than 1000 material presets used in common asset-authoring tools. Many presets share similar roughness and metalness while mainly differing in albedo or texture details. With input from art experts, the original MIO dataset grouped these presets into 14 common material categories. MIO++ extends this vocabulary to 30 categories so that materials with visibly different PBR behavior can be represented more explicitly, for example, by separating smooth, matte, and rusty metal. The detailed category statistics of MIO and MIO++ are reported in Table 3. Each of the 30 MIO++ categories has an independent roughness/metallic pair; the complete values are provided in the Supplementary Materials.

4.2. Data Collection and Annotation

To reduce the domain gap between ordinary 2D images and renderings of 3D assets, we collect and construct image samples under three guidelines: (a) each image contains one dominant foreground object; (b) the collection includes both real-world images and 3D asset renderings; and (c) the images cover diverse camera angles, including less common top and lower-side views. These guidelines make the training distribution closer to the multi-view renderings used by MaterialSeg3D at inference time.
The sources of the collected images are freely accessible public datasets [72,73] and 2D image renderings from 3D objects in website photo libraries [74]. In addition, we also procured some well-designed 3D assets that are used for game development and expanded the data collection by rendering multi-view images of these high-quality assets.
An important feature of MIO++ is the explicit association between semantic material labels and PBR parameters. The material vocabulary and the corresponding roughness/metallic assignments were discussed with professional 3D asset modelers. The modelers examined candidate material spheres drawn from more than 1000 presets in public material libraries and asset-authoring tools, including Adobe Substance 3D Painter 2023, and used their modeling experience to select representative PBR settings. The original MIO label space contains 14 categories. MIO++ expands this space to 30 fine-grained categories, and each MIO++ category has its own independent roughness and metallic values. The complete 30-class parameter table is provided in the Supplementary Materials. The mapping from these 30 categories back to the original 14 MIO categories is used only for the backward-compatible evaluation against the original MIO label space; it is not used in the actual MaterialSeg3D material generation pipeline.
After defining the label space, a specialized annotation team performed pixel-wise dense annotation of the collected images. Only the dominant foreground object was assigned material labels, while the background was treated as a single background class. Annotation was initialized with a Segment Anything [49]-assisted tool and then manually refined. The annotated samples were subsequently passed to other annotators for multiple rounds of checking and correction rather than being accepted from the first pass. Ambiguous boundaries or material regions were handled within this manual review process; the released annotation format does not contain a separate ambiguity label. The annotation process was deliberately paced at approximately 50 images per annotator per day to allow manual boundary and material-label inspection. A numerical inter-annotator agreement score is not reported in the present study; annotation quality was controlled through the SAM-assisted initialization, manual refinement, and repeated checking and correction by different annotators. Figure 5 shows an example of material annotation and PBR alignment, and Figure 6 illustrates the range of object types and viewpoints in MIO++.

4.3. Dataset Distribution

The original Materialized Individual Objects (MIO) dataset contains 23,062 single-object images captured or rendered from diverse viewpoints, with 14 material classes and five object themes: furniture, cars, buildings, musical instruments, and plants. The real/rendered image statistics are reported in Table 4. Approximately 4000 top-view images are included to provide viewpoints that are uncommon in conventional scene datasets; examples from the car theme are shown in Figure 7. MIO++ expands the collection to 115,542 images, 12 object themes, and 30 material types, as summarized in Table 5. Several broad MIO categories are split according to PBR-relevant appearance. For example, metal is divided into smooth metal, matte metal, and rusty metal; wood, fabric, and plastic are also separated into finer subclasses, while new categories such as jade and plaster are added. For the segmentation experiments in this paper, model selection and the detailed results in Table 6 are reported on a held-out MIO++ validation partition. For rendered multi-view samples, all views generated from the same 3D asset are grouped before partitioning and assigned to the same split; therefore, different views of one 3D asset do not cross the training and validation partitions. The independent 3D evaluation is conducted separately on Objaverse-1.0 assets.

5. Method

5.1. Material Segmentation

We formulate 2D material prediction as semantic segmentation tailored to single-object images observed from diverse camera poses. Given an image I, let x i denote the RGB value and y i denote the annotated material label at pixel i. The network encodes the image and predicts a per-pixel probability vector P i = ( p ( i , 1 ) , p ( i , 2 ) , , p ( i , n ) ) over n classes. The final label is obtained by taking the class with maximum predicted probability.
We notice that the Segment Anything Model (SAM) [49] has shown its ability to handle semantic region segmentation on single-object images in previous work [75]. Thus, we formulate the material segmentation network with a modified ViT [76] backbone using pretrained segmentation weights from the SAM-B model. The decode head follows the UPerNet [77] setting with cross-entropy loss as supervision. To prevent possible long-tail problems caused by imbalanced training data, we adopt a class-balanced sampling strategy [78] to enhance the robustness and generalization ability of the model. During training, the cross-entropy loss is calculated as
L = 1 H W i = 1 H W c = 0 n 1 y ( i , c ) log ( p ( i , c ) ) ,
where H and W denote the height and width of the input image, n is the number of classes, and y ( i , c ) and p ( i , c ) denote the ground-truth indicator and predicted probability, respectively, for class c at pixel i.

5.2. MaterialSeg3D

In this section, we introduce MaterialSeg3D, a workflow for generating material information for 3D assets. As shown in Figure 8, the workflow contains three main components: multi-view rendering, material prediction, and material UV generation. First, the input asset is rendered from camera poses covering the object surface. Second, the material segmentation model trained on MIO++ predicts a material label for each visible pixel in every view. Third, the view-level predictions are projected into UV space and combined by weighted voting and region unification. The final material-label UV map is then converted to roughness and metallic maps using the class-specific PBR assignments defined for the 30 MIO++ categories.
Multi-View Rendering. To provide material predictions over the object surface, the rendering cameras are distributed around the full 360 ° azimuth range. We first define five high-value views with elevation/azimuth pairs ( 90 ° , 0 ° ) , ( 15 ° , 0 ° ) , ( 15 ° , 90 ° ) , ( 15 ° , 180 ° ) , and ( 15 ° , 270 ° ) . These views provide one top view and four common inspection views. We then divide the full azimuth range into 12 directions. At each direction, three elevation settings are used: 0 ° and two additional elevations randomly sampled within the positive and negative 30 ° ranges, respectively. An example of the multi-view rendering process is shown in Figure 9. This camera set provides complementary observations over the asset, including top and oblique surfaces. The five predefined views receive a larger contribution during UV voting, as described below. These camera and weighting settings are fixed empirical implementation choices in the current pipeline rather than learned parameters.
Material Prediction. Following Section 5.1, the trained segmentation model predicts a material-label map for each rendered view. These view-level label maps are subsequently projected to UV space.
Material UV Generation. For each rendered view, the predicted material labels are projected to the corresponding texels of the albedo UV map using the known camera and mesh geometry, producing a set of temporary label maps M v i e w = { M 1 , , M n } . A UV texel may be observed by multiple views, so its final label is determined from all valid projected labels rather than by sequentially overwriting the map.
Instead of sequentially updating the material-label UV [79], we use a weighted majority vote. Let m i ( u ) denote the valid material label projected from view i to texel u, and let 1 [ · ] denote the indicator function. The vote score for material class c is
S c ( u ) = i = 1 n w i 1 [ m i ( u ) = c ] , M m a t e r i a l ( u ) = arg max c   S c ( u ) ,
where w i = α for the five predefined high-value views and w i = 1 for the remaining views. We set α = 2 in all experiments; in implementation, this is equivalent to counting each valid prediction from a high-value view twice and each prediction from another view once. Labels outside the valid material-label range are ignored. If no rendered view observes a texel, it remains label 0 (unassigned/background). If two or more valid classes have exactly the same maximum vote count, the implementation selects the first maximum in label-index order, i.e., the smallest material-label ID among the tied classes. The value α = 2 is an empirical setting used consistently throughout the experiments rather than a theoretically optimized parameter.
Region Unification. Some fine-grained labels can fluctuate locally after voting, especially between visually similar subclasses such as smooth and matte metal. To suppress small isolated label changes, we apply a region-level post-processing step. A 2 × 2 erosion is first used to break weak connections, after which connected non-zero regions are identified. Within a connected region, the frequency of material labels is counted. For subclasses belonging to the same meta-material, the majority subclass is used to unify inconsistent local predictions; small minority label regions occupying less than 5 % of the connected component are also replaced by the dominant label. The 5 % threshold, like the view weight, is a fixed empirical setting in this study.
After obtaining the final material-label UV map, each labeled texel is assigned the roughness and metallic values associated with its MIO++ class. The complete 30-class mapping is reported in the Supplementary Materials.

5.3. Plane Simplification on AI-Generated Contents

In addition to human-crafted assets, we apply MaterialSeg3D to assets produced by single-image-to-3D methods [12,13,14]. These generated assets provide a geometry mesh and albedo UV map, but locally irregular triangles and unstable normals can produce unwanted reflection changes on surfaces that should be approximately planar, as illustrated in Figure 10. We therefore use plane-oriented simplification as an auxiliary preprocessing step before material assignment. Related work has explicitly studied structured planar reconstruction for improving geometric consistency and mesh topology in planar regions [80]; our component is instead intended only as lightweight preprocessing for reducing obvious planar artifacts before material assignment. It is not presented as a general mesh-quality benchmark or as a replacement for dedicated geometry-reconstruction methods.
Specifically, as shown in Figure 8, we use four rendered views of the target 3D object (the four high-value side views, excluding the top view) and apply an off-the-shelf 2D surface-normal estimator [81]. The RGB renderings and predicted normal maps are concatenated as cues for K-means clustering of pixels that are likely to belong to the same planar region. The clustered pixels are projected back to 3D vertices through the known camera parameters, after which RANSAC [82] removes spatial outliers. The resulting vertex groups provide plane candidates that guide a mesh-simplification tool. Figure 11 visualizes the stages of this auxiliary preprocessing pipeline. In the present work, its role is assessed qualitatively through the resulting surface and rendering examples; a separate geometry benchmark is discussed as future work in Section 6.4.

6. Experiments

6.1. Implementations and Evaluations

Learning reliable 2D material priors is the first stage of the MaterialSeg3D pipeline. We use a ViT backbone initialized from SAM-B [49] segmentation weights. The optimizer is AdamW [83], with learning rate 6 × 10 5 and weight decay 1 × 10 2 . The batch size is 8, training lasts 80k iterations, and images are resized to 1024 × 1024 . All segmentation experiments are implemented with MMSegmentation [84] and run on four 80GB NVIDIA A100 GPUs. Detailed class- and theme-wise results on the held-out MIO++ validation partition are reported in Table 6. Class-balanced sampling and SAM initialization are kept fixed throughout the reported experiments; they are implementation components rather than separately optimized variables in this study.
Table 6 is intended to expose the class-wise behavior within each object theme rather than only reporting a single aggregate score. The lower values are not confined to one theme: they occur frequently for fine-grained subclasses whose visual difference is subtle, such as smooth/matte/rusty metal, and in material–theme combinations with relatively limited support. This pattern is consistent with the long-tailed occurrence statistics in Table 3; for example, some newly added categories occur much less frequently than common materials such as metal or glass. The theme-wise mIoU in the final row is therefore reported together with the individual class IoUs. A full pixel-level confusion-matrix analysis would provide a more detailed account of subclass confusions, but it is outside the evaluation retained for the current experiments and is left for a dedicated diagnostic study.

6.2. Compared with Previous Work

Material Segmentation. To evaluate the 2D material segmentation component described in Section 5.1, we compare it with five commonly used semantic segmentation backbones on the MIO label space. Table 7 reports mIoU on the MIO validation setting. We also evaluate the model trained on MIO++ under the same 14-category label space. For this comparison only, the 30 MIO++ predictions are mapped back to their MIO meta-materials (for example, smooth, matte, and rusty metal are counted as metal; smooth and rough wood are counted as wood). This evaluation-only mapping makes the MIO and MIO++ models directly comparable and yields an mIoU increase from 83.92% to 86.75%. In the subsequent 3D material generation pipeline, no such collapse is performed: all 30 MIO++ categories and their independent PBR values are retained.
Overall Performance. Figure 12 provides qualitative comparisons in three settings: complete single-image-to-3D generation, texture/material generation on a given mesh, and public human-crafted 3D assets. Wonder3D [12], TripoSR [13], and OpenLRM [14] reconstruct both geometry and appearance from a reference image, whereas MaterialSeg3D assumes that a mesh and albedo UV map are already available. These are therefore different input settings; the single-image-to-3D results are included as end-to-end contextual references and are not intended as a controlled material-only comparison. For methods that can operate on a provided mesh, we additionally show results from Fantasia3D [53], Text2Tex [79], and the online Meshy function (https://app.meshy.ai/). For Fantasia3D, only its appearance-modeling stage is used. We also show materialized renderings for public assets from Tripo3D (https://www.tripo3d.ai/app/ (accessed on 20 August 2026)) and TurboSquid (https://www.turbosquid.com/). These visual comparisons illustrate the different appearance behavior under relighting; Figure 13 provides additional views of assets processed by MaterialSeg3D.
Table 8 further separates the two input settings quantitatively. We select 20 Objaverse-1.0 assets [60] with material information and render their reference and novel views as targets under a common fixed lighting setup. For each asset, three novel camera views are sampled at random while keeping the elevation within a valid viewing range. In the reconstructed input setting, Wonder3D, TripoSR, and OpenLRM reconstruct geometry and appearance from the single reference image; their outputs are resized to the scale of the target asset and aligned in pose before evaluation. These rows measure end-to-end reconstruction quality and should not be interpreted as isolating material prediction. In the predefined mesh + albedo setting, both “Baseline” and “Baseline + Ours” use the same Objaverse mesh and albedo UV. Baseline removes the original material information, whereas Baseline + Ours adds the PBR material maps predicted by MaterialSeg3D. This controlled pair isolates the contribution of the material assignment stage while keeping geometry and albedo fixed. CLIP similarity [30], PSNR, and SSIM are computed against the corresponding target renderings and averaged over the evaluated assets/views.

6.3. Weighted Voting and Region Unification

Appendix A and Figure A1 visualize how multi-view predictions are combined in UV space. A material region that is misclassified in one difficult view can still receive the correct label when the same surface is observed consistently from other views. The bottom examples also show the role of region unification: small local subclass shifts can remain after voting, and the post-processing step suppresses these isolated changes before the class-specific roughness and metallic values are assigned. The current paper uses the camera set, α = 2 , and the 5 % region threshold as fixed pipeline settings. A systematic sensitivity study over these parameters would be useful for characterizing robustness, but is not required to apply the method and is left for future work, as discussed below.

6.4. Limitations and Open Issues

Several limitations remain in the current formulation. First, MaterialSeg3D predicts a discrete material class and then assigns a representative roughness/metallic pair to that class. This design is practical for semantic materialization, but it cannot represent continuous intra-class variation or high-frequency spatial detail. Two wooden regions, for example, may share the same semantic label while having different finish, wear, or local roughness. Future work could refine the class-level maps using albedo-conditioned spatial cues or replace the lookup stage with direct continuous PBR regression when suitable supervision is available.
Second, the quantitative experiments emphasize 2D material segmentation and rendered appearance rather than direct pixel-wise errors between predicted and ground-truth roughness/metalness maps. The Objaverse evaluation in Table 8 controls geometry and albedo in the Baseline/Baseline + Ours comparison and uses target materialized renderings, but authored material maps from heterogeneous public assets are not a calibrated continuous-material benchmark. A dedicated evaluation set with standardized ground-truth PBR maps would enable more direct measurements of roughness/metalness error, UV-space material accuracy, cross-view consistency, and final rendering quality. We consider this an important extension of the present evaluation rather than a claim already established by the current experiments.
Third, the multi-view camera configuration, the high-value-view weight α = 2 , the 5 % region-unification threshold, class-balanced sampling, and SAM initialization are fixed implementation choices. The MIO-versus-MIO++ comparison in Table 7 evaluates the effect of the enlarged/finer dataset, but the present study does not perform a full factorial ablation over all pipeline hyperparameters. Such a study, including the number and elevation of views and alternative vote weights, would provide a more complete sensitivity analysis.
Fourth, the plane-simplification stage is introduced as auxiliary preprocessing for visibly irregular AI-generated meshes and is currently supported mainly by qualitative examples. A dedicated geometry study could quantify normal consistency, surface smoothness, geometric deviation, detail preservation, and rendering changes before and after simplification, and could separately ablate normal estimation, K-means clustering, RANSAC filtering, and mesh simplification. These measurements require a geometry-focused benchmark and are outside the material prediction evaluation used here.
Finally, a reverse-rendering ambiguity remains for AI-generated RGB textures. As shown in Figure 14, illumination can be baked into the generated appearance, making it difficult to determine whether a color variation belongs to the intrinsic albedo or to lighting. Estimating a clean albedo together with material properties under unknown illumination remains an important direction for future work.

7. Conclusions

In this paper, we formulate surface material assignment for 3D assets as material prediction in 2D views followed by fusion in UV space. We expand the original MIO dataset to MIO++, containing 115,542 images, 12 object themes, and 30 fine-grained material categories with independent class-level roughness and metallic assignments. Based on these 2D priors, MaterialSeg3D projects multi-view predictions to UV space, combines them through weighted voting and region unification, and converts the resulting labels to PBR material maps. The expanded MIO++ training data improve material segmentation on the backward-compatible MIO evaluation, while controlled and qualitative 3D experiments show that the assigned material maps can improve relightable appearance when geometry and albedo are available. We also extend the workflow to AI-generated meshes with an auxiliary plane-oriented preprocessing stage. The current class-level PBR mapping and fixed fusion settings favor a practical, lightweight workflow; continuous spatial material prediction, systematic sensitivity analysis, and direct evaluation against standardized PBR ground truth remain useful directions for future work.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/electronics15173885/s1.

Author Contributions

Conceptualization, J.P. and Z.Z.; methodology, J.P., R.G. and Z.Z.; software, R.G. and Z.L.; validation, R.G., S.S. and Z.L.; formal analysis, J.P., R.G. and Z.Z.; investigation, J.P., R.G., Z.Z., S.S. and Y.L.; resources, J.P. and Z.Z.; data curation, R.G., Z.Z., S.S. and Z.L.; writing—original draft preparation, J.P. and R.G.; writing—review and editing, J.P., R.G., Z.Z., S.S. and Y.L.; visualization, R.G., Z.Z., S.S. and Z.L.; supervision, J.P. and Z.Z.; project administration, J.P.; funding acquisition, J.P. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Guangdong Basic and Applied Basic Research Foundation (Grant No. 2024A1515110065).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data and source code that support the findings of this study are openly available at https://github.com/PROPHETE-pro/MaterialSeg3D (accessed on 20 August 2026).

Conflicts of Interest

Authors Zongxing Li and Ziwei Zhu were employed by Dawa Future (Beijing) Imaging Technology Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Appendix A. Additional Visualization of Weighted Voting and Region Unification

Figure A1. Visualization of the segmentation results on multi-view rendering of 3D assets and the colored material UV map acquired from the weighted voting mechanism and region unification.
Figure A1. Visualization of the segmentation results on multi-view rendering of 3D assets and the colored material UV map acquired from the weighted voting mechanism and region unification.
Electronics 15 03885 g0a1

References

  1. Hu, X.; Wang, Y.; Fan, L.; Fan, J.; Peng, J.; Lei, Z.; Li, Q.; Zhang, Z. Semantic anything in 3d gaussians. arXiv 2024, arXiv:2401.17857. [Google Scholar]
  2. Luo, S.; Hu, W. Diffusion probabilistic models for 3d point-cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 2837–2845. [Google Scholar]
  3. Achlioptas, P.; Diamanti, O.; Mitliagkas, I.; Guibas, L. Learning representations and generative models for 3d point clouds. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; pp. 40–49. [Google Scholar]
  4. Yang, G.; Huang, X.; Hao, Z.; Liu, M.Y.; Belongie, S.; Hariharan, B. Pointflow: 3d point-cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4541–4550. [Google Scholar]
  5. Zhou, L.; Du, Y.; Wu, J. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 5826–5835. [Google Scholar]
  6. Mescheder, L.; Oechsle, M.; Niemeyer, M.; Nowozin, S.; Geiger, A. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 4460–4470. [Google Scholar]
  7. Chen, Z.; Zhang, H. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 5939–5948. [Google Scholar]
  8. Zhang, S.H.; Guo, Y.C.; Gu, Q.W. Sketch2model: View-aware 3d modeling from single free-hand sketches. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 6012–6021. [Google Scholar]
  9. Qian, G.; Mai, J.; Hamdi, A.; Ren, J.; Siarohin, A.; Li, B.; Lee, H.Y.; Skorokhodov, I.; Wonka, P.; Tulyakov, S.; et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors. arXiv 2023, arXiv:2306.17843. [Google Scholar]
  10. Gao, J.; Shen, T.; Wang, Z.; Chen, W.; Yin, K.; Li, D.; Litany, O.; Gojcic, Z.; Fidler, S. Get3d: A generative model of high quality 3d textured shapes learned from images. Adv. Neural Inf. Process. Syst. 2022, 35, 31841–31854. [Google Scholar] [CrossRef] [Scilit]
  11. Liu, M.; Xu, C.; Jin, H.; Chen, L.; Xu, Z.; Su, H. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. arXiv 2023, arXiv:2306.16928. [Google Scholar]
  12. Long, X.; Guo, Y.C.; Lin, C.; Liu, Y.; Dou, Z.; Liu, L.; Ma, Y.; Zhang, S.H.; Habermann, M.; Theobalt, C.; et al. Wonder3d: Single image to 3d using cross-domain diffusion. arXiv 2023, arXiv:2310.15008. [Google Scholar]
  13. Tochilkin, D.; Pankratz, D.; Liu, Z.; Huang, Z.; Letts, A.; Li, Y.; Liang, D.; Laforte, C.; Jampani, V.; Cao, Y.P. Triposr: Fast 3d object reconstruction from a single image. arXiv 2024, arXiv:2403.02151. [Google Scholar]
  14. He, Z.; Wang, T. OpenLRM: Open-Source Large Reconstruction Models. Available online: https://github.com/3DTopia/OpenLRM (accessed on 26 May 2026).
  15. Upchurch, P.; Niu, R. A dense material segmentation dataset for indoor and outdoor scene parsing. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 450–466. [Google Scholar]
  16. Bell, S.; Upchurch, P.; Snavely, N.; Bala, K. Material recognition in the wild with the materials in context database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3479–3487. [Google Scholar]
  17. Li, Z.; Gan, R.; Luo, C.; Wang, Y.; Liu, J.; Zhu, Z.; Li, Q.; Yin, X.; Zhang, M.; Zhang, Z.; et al. Materialseg3d: Segmenting dense materials from 2d priors for 3d assets. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, 28 October–1 November 2024; pp. 370–379. [Google Scholar]
  18. Wu, J.; Zhang, C.; Xue, T.; Freeman, B.; Tenenbaum, J. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Adv. Neural Inf. Process. Syst. 2016, 29, 82–90. [Google Scholar]
  19. Smith, E.J.; Meger, D. Improved adversarial systems for 3d object generation and reconstruction. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2017; pp. 87–96. [Google Scholar]
  20. Xie, J.; Zheng, Z.; Gao, R.; Wang, W.; Zhu, S.C.; Wu, Y.N. Learning descriptor networks for 3d shape synthesis and analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8629–8638. [Google Scholar]
  21. Gadelha, M.; Maji, S.; Wang, R. 3d shape induction from 2d views of multiple objects. In Proceedings of the 2017 International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2017; pp. 402–411. [Google Scholar]
  22. Henzler, P.; Mitra, N.J.; Ritschel, T. Escaping plato’s cave: 3d shape from adversarial rendering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9984–9993. [Google Scholar]
  23. Lunz, S.; Li, Y.; Fitzgibbon, A.; Kushman, N. Inverse graphics gan: Learning to generate 3d shapes from unstructured 2d data. arXiv 2020, arXiv:2002.12674. [Google Scholar]
  24. Liu, Y.; Luo, C.; Fan, L.; Wang, N.; Peng, J.; Zhang, Z. CityGaussian: Real-Time High-Quality Large-Scale Scene Rendering with Gaussians. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 265–282. [Google Scholar] [CrossRef] [Scilit]
  25. Liu, Y.; Luo, C.; Mao, Z.; Peng, J.; Zhang, Z. CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025. [Google Scholar]
  26. Zhou, M.; Hou, J.; Luo, C.; Wang, Y.; Zhang, Z.; Peng, J. Scenex: Procedural controllable large-scale scene generation via large-language models. arXiv 2024, arXiv:2403.15698. [Google Scholar]
  27. Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Commun. ACM 2022, 65, 99–106. [Google Scholar] [CrossRef] [Scilit]
  28. Jain, A.; Mildenhall, B.; Barron, J.T.; Abbeel, P.; Poole, B. Zero-shot text-guided object generation with dream fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 867–876. [Google Scholar]
  29. Mohammad Khalid, N.; Xie, T.; Belilovsky, E.; Popa, T. CLIP-Mesh: Generating Textured Meshes from Text Using Pretrained Image-Text Models. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers, Daegu, Republic of Korea, 6–9 December 2022; pp. 25:1–25:8. [Google Scholar] [CrossRef] [Scilit]
  30. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
  31. Poole, B.; Jain, A.; Barron, J.T.; Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. arXiv 2022, arXiv:2209.14988. [Google Scholar]
  32. Lin, C.H.; Gao, J.; Tang, L.; Takikawa, T.; Zeng, X.; Huang, X.; Kreis, K.; Fidler, S.; Liu, M.Y.; Lin, T.Y. Magic3d: High-resolution text-to-3d content creation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 300–309. [Google Scholar]
  33. Abdal, R.; Lee, H.Y.; Zhu, P.; Chai, M.; Siarohin, A.; Wonka, P.; Tulyakov, S. 3davatargan: Bridging domains for personalized editable avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 4552–4562. [Google Scholar]
  34. Chan, E.R.; Lin, C.Z.; Chan, M.A.; Nagano, K.; Pan, B.; De Mello, S.; Gallo, O.; Guibas, L.J.; Tremblay, J.; Khamis, S.; et al. Efficient geometry-aware 3D generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16123–16133. [Google Scholar]
  35. Chan, E.R.; Monteiro, M.; Kellnhofer, P.; Wu, J.; Wetzstein, G. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 5799–5809. [Google Scholar]
  36. Gu, J.; Liu, L.; Wang, P.; Theobalt, C. Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis. arXiv 2021, arXiv:2110.08985. [Google Scholar]
  37. Niemeyer, M.; Geiger, A. Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 11453–11464. [Google Scholar]
  38. Or-El, R.; Luo, X.; Shan, M.; Shechtman, E.; Park, J.J.; Kemelmacher-Shlizerman, I. Stylesdf: High-resolution 3d-consistent image and geometry generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 13503–13513. [Google Scholar]
  39. Skorokhodov, I.; Siarohin, A.; Xu, Y.; Ren, J.; Lee, H.Y.; Wonka, P.; Tulyakov, S. 3d generation on imagenet. arXiv 2023, arXiv:2303.01416. [Google Scholar]
  40. Xu, Y.; Chai, M.; Shi, Z.; Peng, S.; Skorokhodov, I.; Siarohin, A.; Yang, C.; Shen, Y.; Lee, H.Y.; Zhou, B.; et al. DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 4402–4412. [Google Scholar]
  41. Asselin, L.P.; Laurendeau, D.; Lalonde, J.F. Deep SVBRDF estimation on real materials. In Proceedings of the 2020 International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2020; pp. 1157–1166. [Google Scholar]
  42. Deschaintre, V.; Lin, Y.; Ghosh, A. Deep polarization imaging for 3D shape and SVBRDF acquisition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 15567–15576. [Google Scholar]
  43. Deschaintre, V.; Aittala, M.; Durand, F.; Drettakis, G.; Bousseau, A. Single-image svbrdf capture with a rendering-aware deep network. ACM Trans. Graph. 2018, 37, 128. [Google Scholar] [CrossRef] [Scilit]
  44. Martin, R.; Roullier, A.; Rouffet, R.; Kaiser, A.; Boubekeur, T. MaterIA: Single Image High-Resolution Material Capture in the Wild. Comput. Graph. Forum 2022, 41, 163–177. [Google Scholar] [CrossRef] [Scilit]
  45. Gao, D.; Li, X.; Dong, Y.; Peers, P.; Xu, K.; Tong, X. Deep Inverse Rendering for High-Resolution SVBRDF Estimation from an Arbitrary Number of Images. ACM Trans. Graph. 2019, 38, 134:1–134:15. [Google Scholar] [CrossRef] [Scilit]
  46. Deschaintre, V.; Drettakis, G.; Bousseau, A. Guided Fine-Tuning for Large-Scale Material Transfer. Comput. Graph. Forum 2020, 39, 91–105. [Google Scholar] [CrossRef] [Scilit]
  47. Li, X.; Dong, Y.; Peers, P.; Tong, X. Modeling surface appearance from a single photograph using self-augmented convolutional neural networks. ACM Trans. Graph. 2017, 36, 45. [Google Scholar] [CrossRef] [Scilit]
  48. Vecchio, G.; Palazzo, S.; Spampinato, C. SurfaceNet: Adversarial SVBRDF Estimation from a Single Image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 12840–12848. [Google Scholar]
  49. Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. arXiv 2023, arXiv:2304.02643. [Google Scholar]
  50. Vecchio, G.; Sortino, R.; Palazzo, S.; Spampinato, C. Matfuse: Controllable material generation with diffusion models. arXiv 2023, arXiv:2308.11408. [Google Scholar]
  51. Sartor, S.; Peers, P. Matfusion: A generative diffusion model for svbrdf capture. In Proceedings of the SIGGRAPH Asia 2023 Conference Papers, Sydney, Australia, 12–15 December 2023; pp. 86:1–86:10. [Google Scholar]
  52. Vecchio, G.; Martin, R.; Roullier, A.; Kaiser, A.; Rouffet, R.; Deschaintre, V.; Boubekeur, T. ControlMat: A Controlled Generative Approach to Material Capture. arXiv 2023, arXiv:2309.01700. [Google Scholar]
  53. Chen, R.; Chen, Y.; Jiao, N.; Jia, K. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. arXiv 2023, arXiv:2303.13873. [Google Scholar]
  54. Yeh, Y.Y.; Li, Z.; Hold-Geoffroy, Y.; Zhu, R.; Xu, Z.; Hašan, M.; Sunkavalli, K.; Chandraker, M. Photoscene: Photorealistic material and lighting transfer for indoor scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18562–18571. [Google Scholar]
  55. Yuan, L.; Yan, D.; Saito, S.; Fujishiro, I. DiffMat: Latent diffusion models for image-guided material generation. Vis. Inform. 2024, 8, 6–14. [Google Scholar] [CrossRef] [Scilit]
  56. Lopes, I.; Pizzati, F.; de Charette, R. Material Palette: Extraction of Materials from a Single Image. arXiv 2023, arXiv:2311.17060. [Google Scholar]
  57. Ceylan, D.; Deschaintre, V.; Groueix, T.; Martin, R.; Huang, C.H.; Rouffet, R.; Kim, V.; Lassagne, G. MatAtlas: Text-driven Consistent Geometry Texturing and Material Assignment. arXiv 2024, arXiv:2404.02899. [Google Scholar]
  58. Zhang, Y.; Tu, Z.; Lian, W.; Hu, Y.; Sun, S.; Xiao, Y.; Cheng, Y. BANet: Bidirectional Feature Aggregation and Adaptive Multi-Scene Perception-Based Lane Detection for Autonomous Driving. IEEE Trans. Intell. Transp. Syst. 2026, 27, 9713–9723. [Google Scholar] [CrossRef] [Scilit]
  59. Wang, Y.; Peng, J.; Zhang, G.; Luo, C.; Xu, S.; Zhang, M.; Zhang, Z. FurniScene: A Large-Scale 3D Room Dataset with Intricate Furnishing Scenes. Int. J. Comput. Vis. 2026, 134, 125. [Google Scholar] [CrossRef] [Scilit]
  60. Deitke, M.; Schwenk, D.; Salvador, J.; Weihs, L.; Michel, O.; VanderBilt, E.; Schmidt, L.; Ehsani, K.; Kembhavi, A.; Farhadi, A. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 13142–13153. [Google Scholar]
  61. Deitke, M.; Liu, R.; Wallingford, M.; Ngo, H.; Michel, O.; Kusupati, A.; Fan, A.; Laforte, C.; Voleti, V.; Gadre, S.Y.; et al. Objaverse-xl: A universe of 10m+ 3d objects. arXiv 2023, arXiv:2307.05663. [Google Scholar]
  62. Kasper, A.; Xue, Z.; Dillmann, R. The kit object models database: An object model database for object recognition, localization and manipulation in service robotics. Int. J. Robot. Res. 2012, 31, 927–934. [Google Scholar] [CrossRef] [Scilit]
  63. Calli, B.; Walsman, A.; Singh, A.; Srinivasa, S.; Abbeel, P.; Dollar, A.M. Benchmarking in manipulation research: The ycb object and model set and benchmarking protocols. arXiv 2015, arXiv:1502.03143. [Google Scholar]
  64. Singh, A.; Sha, J.; Narayan, K.S.; Achim, T.; Abbeel, P. Bigbird: A large-scale 3d database of object instances. In Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2014; pp. 509–516. [Google Scholar]
  65. Sun, X.; Wu, J.; Zhang, X.; Zhang, Z.; Zhang, C.; Xue, T.; Tenenbaum, J.B.; Freeman, W.T. Pix3d: Dataset and methods for single-image 3d shape modeling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 2974–2983. [Google Scholar]
  66. Downs, L.; Francis, A.; Koenig, N.; Kinman, B.; Hickman, R.; Reymann, K.; McHugh, T.B.; Vanhoucke, V. Google scanned objects: A high-quality dataset of 3d scanned household items. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2022; pp. 2553–2560. [Google Scholar]
  67. Park, K.; Rematas, K.; Farhadi, A.; Seitz, S.M. Photoshape: Photorealistic materials for large-scale shape collections. arXiv 2018, arXiv:1809.09761. [Google Scholar]
  68. Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; Xiao, J. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 1912–1920. [Google Scholar]
  69. Wu, R.; Xiao, C.; Zheng, C. Deepcad: A deep generative network for computer-aided design models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 6772–6782. [Google Scholar]
  70. Koch, S.; Matveev, A.; Jiang, Z.; Williams, F.; Artemov, A.; Burnaev, E.; Alexa, M.; Zorin, D.; Panozzo, D. ABC: A Big CAD Model Dataset for Geometric Deep Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 9601–9611. [Google Scholar]
  71. Bell, S.; Upchurch, P.; Snavely, N.; Bala, K. OpenSurfaces: A richly annotated catalog of surface appearance. ACM Trans. Graph. 2013, 32, 111. [Google Scholar]
  72. images.cv. CV Image Dataset. 2024. Available online: https://images.cv (accessed on 26 May 2026).
  73. Yang, L.; Luo, P.; Loy, C.C.; Tang, X. A large-scale car dataset for fine-grained categorization and verification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3973–3981. [Google Scholar]
  74. Aubry, M.; Maturana, D.; Efros, A.A.; Russell, B.C.; Sivic, J. Seeing 3d chairs: Exemplar part-based 2d-3d alignment using a large dataset of cad models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 3762–3769. [Google Scholar]
  75. Wu, T.; Li, Z.; Yang, S.; Zhang, P.; Pan, X.; Wang, J.; Lin, D.; Liu, Z. HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image. In Proceedings of the SIGGRAPH Asia 2023 Conference Papers, Sydney, Australia, 12–15 December 2023; pp. 53:1–53:10. [Google Scholar]
  76. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
  77. Xiao, T.; Liu, Y.; Zhou, B.; Jiang, Y.; Sun, J. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 418–434. [Google Scholar]
  78. Contributors, M. MMEngine: OpenMMLab Foundational Library for Training Deep Learning Models. Available online: https://github.com/open-mmlab/mmengine (accessed on 26 May 2026).
  79. Chen, D.Z.; Siddiqui, Y.; Lee, H.Y.; Tulyakov, S.; Nießner, M. Text2tex: Text-driven texture synthesis via diffusion models. arXiv 2023, arXiv:2303.11396. [Google Scholar]
  80. Gan, R.; Peng, J.; Liu, Y.; Luo, C.; Li, Q.; Zhang, Z. GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation. arXiv 2025, arXiv:2510.17095. [Google Scholar]
  81. Yin, W.; Zhang, C.; Chen, H.; Cai, Z.; Yu, G.; Wang, K.; Chen, X.; Shen, C. Metric3d: Towards zero-shot metric 3d prediction from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 9043–9053. [Google Scholar]
  82. Derpanis, K.G. Overview of the RANSAC Algorithm. Image 2010, 4, 2–3. [Google Scholar]
  83. Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. arXiv 2018, arXiv:1711.05101. [Google Scholar]
  84. Contributors, M. MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark. 2020. Available online: https://github.com/open-mmlab/mmsegmentation (accessed on 1 June 2026).
  85. Liu, Z.; Mao, H.; Wu, C.Y.; Feichtenhofer, C.; Darrell, T.; Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 11976–11986. [Google Scholar]
  86. Sun, K.; Zhao, Y.; Jiang, B.; Cheng, T.; Xiao, B.; Liu, D.; Mu, Y.; Wang, X.; Liu, W.; Wang, J. High-resolution representations for labeling pixels and regions. arXiv 2019, arXiv:1904.04514. [Google Scholar]
  87. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
  88. He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16000–16009. [Google Scholar]
Figure 1. (a) Renderings of raw 3D assets with only Albedo information. (b) Renderings of processed assets with material information under different lighting conditions.
Figure 1. (a) Renderings of raw 3D assets with only Albedo information. (b) Renderings of processed assets with material information under different lighting conditions.
Electronics 15 03885 g001
Figure 2. Comparison of 3D assets rendered with and without PBR material information under the same lighting conditions.
Figure 2. Comparison of 3D assets rendered with and without PBR material information under the same lighting conditions.
Electronics 15 03885 g002
Figure 3. Examples of material assignment issues on AI-generated assets. (a) A nearly uniform PBR setting is applied to object parts with different material semantics. (b) Material labels vary within a region that should have a consistent semantic material.
Figure 3. Examples of material assignment issues on AI-generated assets. (a) A nearly uniform PBR setting is applied to object parts with different material semantics. (b) Material labels vary within a region that should have a consistent semantic material.
Electronics 15 03885 g003
Figure 4. Rendering results of 3D assets assigned different roughness values. Lower roughness produces sharper specular highlights, whereas higher roughness broadens and attenuates the specular response.
Figure 4. Rendering results of 3D assets assigned different roughness values. Lower roughness produces sharper specular highlights, whereas higher roughness broadens and attenuates the specular response.
Electronics 15 03885 g004
Figure 5. Single case of the material class annotation and the mapping with PBR material spheres.
Figure 5. Single case of the material class annotation and the mapping with PBR material spheres.
Electronics 15 03885 g005
Figure 6. Examples of data and annotations in the MIO++ dataset, covering diverse object categories and camera viewpoints.
Figure 6. Examples of data and annotations in the MIO++ dataset, covering diverse object categories and camera viewpoints.
Electronics 15 03885 g006
Figure 7. Examples from the Cars meta-class in the MIO dataset.
Figure 7. Examples from the Cars meta-class in the MIO dataset.
Electronics 15 03885 g007
Figure 8. Overall framework of MaterialSeg3D. The material segmentation model is trained on the MIO++ dataset before inference. For human-crafted assets, MaterialSeg3D predicts material labels from multi-view renderings and generates PBR material maps. For AI-generated assets, an auxiliary plane-simplification stage is applied to visibly poor mesh topology before the same material workflow.
Figure 8. Overall framework of MaterialSeg3D. The material segmentation model is trained on the MIO++ dataset before inference. For human-crafted assets, MaterialSeg3D predicts material labels from multi-view renderings and generates PBR material maps. For AI-generated assets, an auxiliary plane-simplification stage is applied to visibly poor mesh topology before the same material workflow.
Electronics 15 03885 g008
Figure 9. Example of multi-view rendering for material prediction. Five high-value camera views use elevation/azimuth pairs of ( 90 ° , 0 ° ) , ( 15 ° , 0 ° ) , ( 15 ° , 90 ° ) , ( 15 ° , 180 ° ) , and ( 15 ° , 270 ° ) . The full 360 ° azimuth range is then divided into 12 directions, each using three elevation settings: 0 ° and two additional elevations sampled within the positive and negative 30 ° ranges. Together, these views provide coverage of the object surface, including top and oblique regions.
Figure 9. Example of multi-view rendering for material prediction. Five high-value camera views use elevation/azimuth pairs of ( 90 ° , 0 ° ) , ( 15 ° , 0 ° ) , ( 15 ° , 90 ° ) , ( 15 ° , 180 ° ) , and ( 15 ° , 270 ° ) . The full 360 ° azimuth range is then divided into 12 directions, each using three elevation settings: 0 ° and two additional elevations sampled within the positive and negative 30 ° ranges. Together, these views provide coverage of the object surface, including top and oblique regions.
Electronics 15 03885 g009
Figure 10. Rendering examples on poor-quality meshes. The top example is generated by Wonder3D [12], and the bottom example is generated by TripoSR [13].
Figure 10. Rendering examples on poor-quality meshes. The top example is generated by Wonder3D [12], and the bottom example is generated by TripoSR [13].
Electronics 15 03885 g010
Figure 11. Visualization of the plane-simplification process at each stage.
Figure 11. Visualization of the plane-simplification process at each stage.
Electronics 15 03885 g011
Figure 12. Detailed visual comparisons in three settings: single-image-to-3D generation, texture/material generation on a given mesh, and human-crafted public 3D assets rendered under different lighting conditions.
Figure 12. Detailed visual comparisons in three settings: single-image-to-3D generation, texture/material generation on a given mesh, and human-crafted public 3D assets rendered under different lighting conditions.
Electronics 15 03885 g012
Figure 13. Rendering results after applying the PBR material predicted by MaterialSeg3D to the colored meshes.
Figure 13. Rendering results after applying the PBR material predicted by MaterialSeg3D to the colored meshes.
Electronics 15 03885 g013
Figure 14. Visualization of baked illumination in a 3D mesh generated by a single-image-to-3D method. The example is produced with Wonder3D [12].
Figure 14. Visualization of baked illumination in a 3D mesh generated by a single-image-to-3D method. The example is produced with Wonder3D [12].
Electronics 15 03885 g014
Table 1. Summary of the main extensions from the ACM MM 2024 conference version to this journal version.
Table 1. Summary of the main extensions from the ACM MM 2024 conference version to this journal version.
AspectACM MM 2024 VersionThis Journal Version
DatasetMIO: 23,062 imagesMIO++: 115,542 images
Object themes512
Material-label space14 categories30 fine-grained categories
PBR assignment14 class-level settings30 independent class-level settings
AI-generated meshesOriginal material workflowAdded plane-oriented preprocessing
EvaluationOriginal MIO evaluationMIO++ analysis and expanded 3D evaluations
Table 2. Aggregated material selection counts from an exploratory survey of 100 participants, each evaluating 10 indoor/outdoor scene images (1000 image–participant observations in total).
Table 2. Aggregated material selection counts from an exploratory survey of 100 participants, each evaluating 10 indoor/outdoor scene images (1000 image–participant observations in total).
Material LabelNumberMaterial LabelNumber
metal935brick186
wood842porcelain163
plastic768clay terracotta154
glass712concrete152
paint626nylon75
rubber524rusty metal53
leather437stone46
fabric391bone25
fruit&leaf273bamboo22
flower252others181
Table 3. Occurrence statistics for material categories in MIO and MIO++. In MIO, ceramic (No.18) is named porcelain. The ↑ symbol indicates MIO++ subclasses whose counts are combined when mapped back to an original MIO meta-material.
Table 3. Occurrence statistics for material categories in MIO and MIO++. In MIO, ceramic (No.18) is named porcelain. The ↑ symbol indicates MIO++ subclasses whose counts are combined when mapped back to an original MIO meta-material.
Label Id.Material LabelMIOMIO++
1rubber53247244
2smooth metal894618,047
3matte metal15,429
4rusty metal1560
5paint34964003
6glass580219,637
7leather34173496
8smooth wood708811,691
9rough wood6156
10rough fabric337312,270
11smooth fabric1283
12plush fabric2199
13smooth plastic692811,807
14rough plastic10,939
15brick10172689
16concrete7944306
17clay9102290
18ceramic9216604
19flower16772577
20fruit17424758
21leaf3020
22meat-3345
23gel-2438
24cream-1957
25biscuit-5287
26shell-756
27water-1596
28jade-519
29pearl-212
30plaster-2014
Table 4. Number of rendered images and real images of every meta-class in the MIO dataset.
Table 4. Number of rendered images and real images of every meta-class in the MIO dataset.
Classname (Abbr.)Rendered ImageReal ImageTotal
Furniture (fur.)415254559607
Cars (car)193541176052
Buildings (bui.)41817522170
Musical Instrument (ins.)62716372264
Plants (pla.)55224172969
Table 5. Number of images in each object theme in the MIO dataset and the number of images newly added in the MIO++ dataset.
Table 5. Number of images in each object theme in the MIO dataset and the number of images newly added in the MIO++ dataset.
Classname (Abbr.)MIONew in MIO++Total
Furniture (fur.)9607249912,106
Car (car)6052555811,610
Building (bui.)217064888658
Musical Instrument (ins.)2264833710,601
Statue (sta.)-24002400
Household appliance (hou.)-11,99111,991
Kitchen ware (kit.)-11,99211,992
Weapon and tool (wea.)-11,89111,891
Clothing (clo.)-13,90613,906
Food (foo.)-14,97014,970
Jewelry (jew.)-28082808
Plants (pla.)2969-2969
Total23,06292,840115,542
Table 6. Per-class IoU (%) on the MIO++ validation set, stratified by object theme. “–” denotes a material–theme combination for which no valid IoU is reported in this stratified evaluation. The last row reports the theme-wise mIoU. Class occurrence frequencies in the full MIO++ collection are given in Table 3.
Table 6. Per-class IoU (%) on the MIO++ validation set, stratified by object theme. “–” denotes a material–theme combination for which no valid IoU is reported in this stratified evaluation. The last row reports the theme-wise mIoU. Class occurrence frequencies in the full MIO++ collection are given in Table 3.
bui.carclo.hou.foo.jew.kit.mus.sta.fur.wea.pla.
rubber73.5773.848.4659.98
smooth metal55.7841.8141.6711.0670.7852.7376.6723.4119.9142.02
matte metal49.7842.8545.1121.5445.443.5752.4536.6184.63
rusty metal49.9119.731.1731.3375.35
paint79.93
glass76.384.9757.8280.0618.6369.2859.8981.82
leather52.2566.2745.2461.6362.2626.74
smooth wood51.3573.8674.2743.5566.0848.76
rough wood71.2457.355.9448.9622.4961.1354.1878.73
rough fabric64.2288.9744.2943.6624.3684.1867.95
smooth fabric29.48
plush fabric47.5722.34
smooth plastic61.0729.3139.3323.2466.4343.0651.84
rough plastic36.5642.3271.263.8142.4552.7624.39
brick76.73
concrete80.27
clay64.5932.5738.8180.18
ceramic17.8445.7277.2547.3586.48
flower27.4194.19
fruit87.2669.08
leaf76.1239.0457.2955.6336.8692.30
meat67.25
gel53.82
cream72.54
biscuit55.377.0444.71
shell75.57
water43.2137.18
jade61.16
pearl54.77
plaster89.56
mIoU67.6862.5952.5046.4753.7449.0846.6253.1243.0953.8253.58486.38
Table 7. Quantitative results of semantic segmentation methods on the MIO validation label space. Ours(MIO) denotes the model trained on MIO. Ours(MIO++) denotes the model trained on MIO++; only for this backward-compatible comparison, its 30 fine-grained predictions are mapped to the corresponding 14 MIO meta-material categories. The actual MaterialSeg3D pipeline retains all 30 MIO++ categories and their independent PBR assignments.
Table 7. Quantitative results of semantic segmentation methods on the MIO validation label space. Ours(MIO) denotes the model trained on MIO. Ours(MIO++) denotes the model trained on MIO++; only for this backward-compatible comparison, its 30 fine-grained predictions are mapped to the corresponding 14 MIO meta-material categories. The actual MaterialSeg3D pipeline retains all 30 MIO++ categories and their independent PBR assignments.
MethodMIO Dataset (%)
carfur.bui.ins.pla.mIoU
ConvNeXt [85]71.0374.8569.3372.4076.7272.87
HRNet [86]75.7179.9476.3780.1481.3578.70
ViT [76]73.9677.6775.5379.4578.6677.05
Swin-T [87]75.0979.0478.4580.9281.4078.98
MAE [88]76.4282.0677.5982.7485.9280.95
Ours(MIO)81.8385.2281.7684.3986.3883.92
Ours(MIO++)85.1487.3584.4988.8287.9786.75
Bold values indicate the best performance among all compared methods.
Table 8. Quantitative evaluation on 20 Objaverse-1.0 assets. The first three rows use geometry/appearance reconstructed from one reference image and provide end-to-end context. The last two rows use the same predefined mesh and albedo UV; their difference isolates the effect of adding MaterialSeg3D-predicted PBR material. Metrics are averaged over the evaluated assets/views.
Table 8. Quantitative evaluation on 20 Objaverse-1.0 assets. The first three rows use geometry/appearance reconstructed from one reference image and provide end-to-end context. The last two rows use the same predefined mesh and albedo UV; their difference isolates the effect of adding MaterialSeg3D-predicted PBR material. Metrics are averaged over the evaluated assets/views.
MethodInput SettingCLIP Similarity↑PSNR↑SSIM↑
ReferenceNovelReferenceNovelReferenceNovel
Wonder3D [12]Reconstructed0.850.8416.0615.830.780.75
TripoSR [13]0.930.9016.9316.140.790.76
OpenLRM [14]0.920.8716.3015.370.770.76
BaselinePredefined mesh + albedo0.930.9316.2816.300.790.78
Baseline + Ours0.98 0.9720.7218.390.850.84
The symbol “↑” indicates that higher values represent better performance. Bold text highlights the best performance among the compared methods.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Peng, J.; Gan, R.; Shen, S.; Li, Z.; Liu, Y.; Zhu, Z. MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors. Electronics 2026, 15, 3885. https://doi.org/10.3390/electronics15173885

AMA Style

Peng J, Gan R, Shen S, Li Z, Liu Y, Zhu Z. MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors. Electronics. 2026; 15(17):3885. https://doi.org/10.3390/electronics15173885

Chicago/Turabian Style

Peng, Junran, Ruitong Gan, Silei Shen, Zongxing Li, Yan Liu, and Ziwei Zhu. 2026. "MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors" Electronics 15, no. 17: 3885. https://doi.org/10.3390/electronics15173885

APA Style

Peng, J., Gan, R., Shen, S., Li, Z., Liu, Y., & Zhu, Z. (2026). MaterialSeg3D++: Large-Scale Material Prediction for 3D Assets from 2D Priors. Electronics, 15(17), 3885. https://doi.org/10.3390/electronics15173885

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop