LinxiStart a commission ↗
← Journal / Journal guide

Single-Image 3D Generation: Challenges and Research Directions for Preserving Individual Identity

Linxi Editorial · · 11 min read

Abstract: This article examines static, full-body 3D generation from a single image of a person or pet. It explores how feed-forward reconstruction, generation in a 3D latent space, multi-view inference, and pixel alignment can complement one another, and outlines a potential framework for preserving the subject’s distinctive identity. The discussion focuses on local feature extraction, observation reliability, completion of occluded regions, and consistency between geometry and materials. It also reviews the relevant research foundations and proposes an evaluation strategy. The proposed improvements remain to be validated experimentally.

Keywords: image-to-3D generation; people and pets; identity fidelity; pixel alignment; observation reliability; multi-view consistency

1. Research Goal and Problem Definition

One clear application of image-to-3D generation is to turn a single photograph of a person or pet into a model that users can rotate and inspect in a web browser. Such a model needs to meet three requirements: preserve the subject’s distinctive features from the input viewpoint, maintain a plausible structure from other angles, and keep its geometry and textures consistently aligned.

For people, recognizability often comes down to facial shape, hairstyle, and physique. For pets, ear shape, muzzle proportions, and coat markings can be just as important. These features suggest a research goal of identity-preserving 3D generation: faithfully reconstructing what is visible while completing unseen regions in ways that fit the individual subject.

placeholder — Original figure accompanying the discussion of identity-preserving 3D generation

A single image inevitably leaves gaps in the available information. A photograph captures a projection from just one viewpoint, and the same pixel can correspond to surfaces at different depths in 3D space. The construction of clothing on a person’s back, a pet’s occluded hind legs, or the markings on its far side cannot be determined directly from that photograph. Generating a complete model therefore requires both evidence from the image and 3D priors learned from training data.

Visible and invisible regions consequently call for different objectives. Visible regions demand fidelity; unseen regions demand plausibility. Several different reconstructions of the back of a subject may all be consistent with the same photograph. An algorithm can encourage these possibilities to agree with the known shape and appearance, but it cannot guarantee recovery of details that were never photographed.

2. Existing Methods and Technical Foundations

Existing image-to-3D research offers several complementary building blocks.

Approach

Representative methods

Relevant capabilities

Predict a 3D representation directly from an image

LRM, TripoSR

Map image features into a 3D representation, with silhouette and local rendering supervision to improve detail.

Generate a complete object in a 3D latent space

TRELLIS, TRELLIS.2

Learn distributions of complete shapes and appearances, providing generative priors for unseen regions.

Generate additional views before reconstructing the object

Zero123++, InstantMesh

Infer multiple views from a single image and pass them to a reconstruction module.

Explicitly associate image pixels with 3D locations

Pixal3D

Deliver image conditioning more precisely to the corresponding spatial locations, improving fidelity from the input viewpoint.

These approaches address different aspects of system design and can be used together. Pixel alignment, for example, concerns how image information enters the 3D generation process, while generation in a 3D latent space concerns how a complete object is represented and generated. The two are not inherently mutually exclusive.

placeholder — Original figure accompanying the overview of existing image-to-3D methods

2.1 Feed-Forward Reconstruction: LRM and TripoSR

LRM uses an image encoder and a Transformer to map the input into triplane features, then predicts density and color through spatial queries. TripoSR builds on this approach with a mask loss and local rendering supervision. Local supervision is particularly relevant to people and pets: faces, eyes, and ears often occupy only a small part of a full-body photograph, so training needs to give these regions sufficient attention.

2.2 Generation in a 3D Latent Space: The TRELLIS Family

TRELLIS represents 3D objects using sparse structures and local latent variables, and generates them through flow matching. TRELLIS.2 extends this approach with O-Voxel representations for geometry and materials. These methods draw on priors learned from complete 3D assets to supply structural information that a single image cannot provide directly.

2.3 Multi-View Inference: Zero123++ and InstantMesh

Zero123++ supplements the input by jointly generating several additional views. InstantMesh already combines multi-view diffusion with LRM-based reconstruction. Generated views, however, are still model predictions. If they offer conflicting interpretations of ears, limbs, or markings, the reconstruction process may embed those inconsistencies in the final 3D model.

2.4 Pixel-Aligned Conditioning: Pixal3D

Pixal3D improves fidelity to the input through an explicit projection relationship. In camera coordinates, a 3D location can be projected onto the image so that features are sampled from the corresponding pixel. In words, the mapping is:

Pixel-aligned feature at a 3D location ← the image feature sampled at that location’s camera projection.

Here, x denotes the 3D location, π_K denotes the camera projection, and F(I) is the image feature map. This mapping tells the model where in the image to obtain information for a given spatial location. Multiple points along the same camera ray can still receive identical image features, however, so the model must infer the actual depth and surface location.

placeholder — Original figure accompanying the explanation of pixel-aligned conditioning

3. An Overall Architecture for People and Pets

These capabilities already appear in public work from several teams. Tencent’s Hunyuan and Deemos’s CLAY explore combining shape generation with multi-view material generation. TripoSG focuses on image-conditioned 3D shape generation. Meta’s SAM 3D Objects addresses occlusion and the recovery of complete objects from real photographs, while Meshy T2 explores generation from coarse structures to detailed meshes. These efforts provide useful references for combining techniques, although product results alone do not establish that the systems follow the same internal pipeline.

It is also important to distinguish a method’s original paper from later releases. The original Pixal3D paper focuses primarily on geometry generation, whereas its official updated release is already built on TRELLIS.2 and supports texture generation and GLB export. Research on people and pets can therefore start from an existing system with complete generation capabilities and focus improvements on preserving individual features and making occlusion completion more reliable.

Together, these techniques suggest the following pipeline, which remains to be validated:

Extract full-body and local features → Generate an initial 3D structure → Estimate visibility and observation reliability → Refine geometry and appearance → Evaluate consistency across viewpoints.

Image 4 placeholder — Original figure accompanying the proposed overall pipeline

4. Key Module Designs

4.1 Extracting Full-Body Features and Local Identity Cues

The input stage should retain both global and local information. Full-body features describe pose, body shape, and the outline of clothing or fur. High-resolution local features capture facial details in people and ear shape, muzzle structure, and coat markings in pets. Local crops should preserve their coordinate relationships with the original image so that their features can be mapped to the correct 3D locations. Otherwise, the result may contain sharp local details that are misplaced or incorrectly proportioned.

4.2 Initial Geometry Generation and Visibility Estimation

An existing backbone can generate the initial 3D structure. This provides a hypothesis for the complete shape and can also be used to estimate visible surfaces, depth, and normals from the original viewpoint. These estimates should retain uncertainty: errors in the coarse model can produce incorrect visibility estimates and mislead subsequent refinement.

4.3 Fusing Conditions According to Observation Reliability

Building on the initial geometry, a proposed observation reliability control mechanism would adjust how strongly local image information influences the 3D result. Its decisions would reflect image quality, boundary conditions, and the reliability of geometric correspondence. A conceptual formulation combines:

Global semantic conditioning + local image features mapped into a compatible feature space and weighted by their reliability.

The global component supplies overall semantic context, P maps local features into a compatible feature space, and a reliability term expresses how trustworthy the local observation is. This is an abstract description of a candidate mechanism; its design and effectiveness still require further work and experimental validation.

Clearly visible facial features and markings with trustworthy projection correspondences should receive stronger image constraints. Around fur boundaries, occlusions, and blurred regions, local evidence should exert less force, leaving room for generative priors to complete the model. The original Pixal3D paper notes that segmentation boundary noise can be amplified by back-projection, making this mechanism especially worth investigating around hair strands, fur, ear tips, and tail boundaries.

Image 5 placeholder — Original figure accompanying the discussion of observation reliability and local feature fusion

Reliability also needs a supervision strategy. Renderings of known 3D assets can provide visibility and correspondence information. Perturbations to segmentation, camera parameters, and occlusion can then help the model learn which observations to trust. At the same time, training must ensure that the model continues to respond to reliable evidence in the original image, preventing it from simply suppressing all local weights and relying entirely on its priors.

4.4 Completing Unseen Regions and Constraining Generated Views

Unseen regions can be completed directly using 3D priors or with newly generated views constrained by the coarse geometry. In the latter case, the system should explicitly distinguish actual photographs from generated views and limit the influence of unreliable generated information on features already established by the input. When completing a cat’s side, for example, the model should preserve the ear shape and muzzle proportions visible in the photograph. When completing a person’s back, it should extend the structural cues already present in the visible clothing and hairstyle.

4.5 Maintaining Consistency Between Geometry and Materials

Geometry refinement and appearance generation need to constrain one another. Color patches on a pet’s face should occupy the correct surface regions, and facial textures on a human model should align with the underlying shape. Shadows and highlights in the photograph also need to be distinguished from the material’s base color. Otherwise, rotating the model or changing the lighting in a web viewer may reveal unnatural shading.

4.6 Adapting the Model to People and Pets

People and pets can share some image-encoding and 3D generation capabilities, but they require different structural priors. Human joints, clothing, and hair impose different constraints from the limbs, ears, tails, and fur distributions of cats and dogs. Category conditioning or local adaptation modules could account for these differences, but their ability to handle complex poses and occlusion must be evaluated separately.

Image 6 placeholder — Original figure accompanying the discussion of category-specific adaptation for people and pets

5. Research Novelty and Validation

5.1 Existing Work and Open Questions

The research value of this approach needs to be assessed in light of prior work. Human-LRM already combines coarse human reconstruction, novel-view diffusion, and fine reconstruction. SiTH explores generating a back view before reconstructing a complete person. Unique3D also uses visibility and other information for weighted fusion. Simply connecting existing models, or applying visibility estimates or local supervision on their own, is therefore insufficient to establish algorithmic novelty.

A more specific research question is: For people and pets, can learned, calibrated estimates of local observation reliability jointly guide geometry and appearance refinement to reduce identity drift caused by complex boundaries and occlusion? This question identifies where the proposed improvement would occur and gives it a testable objective.

5.2 Controlled Experiments and Data Splits

The following comparisons can help separate the effects of domain-specific data from those of the proposed algorithmic modules.

Experimental setting

Main change

Purpose

Original baseline

Use an existing model.

Establish its initial generation capabilities.

Domain-specific fine-tuning

Introduce data on people and pets.

Identify gains from adapting to the target domain.

Reliability control

Add the reliability module to the same fine-tuned baseline.

Test whether weighting the conditioning reduces error propagation.

Local identity supervision

Strengthen supervision of key regions on the same fine-tuned baseline.

Test whether local details are better preserved.

Combined approach

Use both proposed modules.

Assess their combined benefits and interactions.

Comparisons should control for training data, inference settings, and evaluation conditions. Training and test sets should be split by subject for both people and pets, so that different views of the same individual or 3D asset do not appear across the two sets.

5.3 Evaluating Quality Across Viewpoints

Evaluation should cover three areas:

Fidelity of visible features: How well the silhouette, facial shape, ear shape, and textures match the original image from its viewpoint.

Plausibility of the 3D structure: Whether side and back views have reasonable shapes and limb relationships, and whether they contain obvious generation errors.

Consistency of appearance across views: Whether markings and materials remain stably aligned with the geometry as the model rotates.

A web preview offers an intuitive final check. As the model turns, does the facial shape change smoothly? Do the ears distort? Do markings stay in the correct places? Do the limbs retain a plausible structure? Building generation and evaluation around these questions makes identity fidelity observable in concrete terms and helps determine which combinations of techniques merit further development.

References and Further Reading

The source article links papers for the core methods in their respective technical sections. It also lists the following resources for further information on the teams and models discussed:

[Pixal3D 官方代码与发布说明](https://github.com/TencentARC/Pixal3D)

[Hunyuan3D 2.5 技术报告](https://arxiv.org/abs/2506.16504)

[TripoSG 论文](https://arxiv.org/abs/2502.06608)

[CLAY 论文](https://arxiv.org/abs/2406.13897)

[Meshy T2 论文](https://arxiv.org/abs/2607.28675)

[SAM 3D Objects 官方说明](https://github.com/facebookresearch/sam-3d-objects)

Made for your story

Give a memory a form.

Start your commission ↗