Guides · 3 min read

What makes a good input image for 3D generation

The difference between a usable mesh and a wasted generation is almost always the photograph. Six things that matter, and three that do not.

28 August 2026

People blame the model. It is usually the photograph.

Image-to-3D reconstruction has to infer a whole object from one flat view. What you give it determines almost everything about what comes back, and the rules are simpler than they look.

Six things that matter

One subject

The single biggest cause of a bad result. If there are three objects in frame, the model has to decide which one you meant. It usually picks the largest and merges the rest into it, which is how you get a chair with a table growing out of its side.

Crop to the thing you want before you upload.

A background it can separate from

The engines remove the background themselves, and they are good at it. They are good at it when there is a background to remove: a clear edge between subject and surroundings.

They struggle when the subject is the same colour as what is behind it, when the lighting flattens the silhouette, or when the object is photographed against a pattern that reads as texture.

A plain wall beats a studio backdrop beats a cluttered desk.

Even lighting

Hard shadows get baked into the texture as though they were paint. A shadow across a face becomes a dark patch on the model that no relighting removes, because it is now part of the surface.

Overcast daylight is close to ideal. So is any large soft source. Direct sun and a phone flash are the two worst options.

Resolution, within reason

More pixels give more texture detail, up to the point where the mesh cannot carry it. Past roughly 2000 pixels on the long edge you are not gaining much.

Below about 500 you are losing a lot.

A three-quarter view

A dead-on front view gives the model nothing about depth. A pure profile gives it nothing about width. A three-quarter angle, turned maybe thirty degrees and slightly above, shows two faces of the object at once, and that is what lets the reconstruction infer the third.

This one change improves results more than any other on this list.

A subject with a readable silhouette

Objects that are mostly negative space (a bicycle, a chain-link fence, a potted fern) reconstruct badly, because the silhouette that carries most of the information is mostly holes.

Solid objects with clear outlines are what this technology is good at.

Three things that do not matter

Camera quality. A recent phone is fine. Sensor size is not the bottleneck; lighting and angle are.

File format. JPG, PNG and WebP are all equivalent for this. Do not convert anything before uploading.

Prompt-style detail in the filename or title. The image is the input. The title is for you.

What to do with a flat drawing

Line art and flat illustrations give poor geometry, because there is no shading for the model to read depth from. Two ways around it:

  • Render or shade it first, so the drawing has light and volume, then feed that in.
  • Describe it instead. Text-to-3D does not need an image at all, and for something that only exists as a sketch, a good written description often beats a bad photograph of the sketch.

When to spend more

A failed generation refunds its tokens here, so the cost of a bad attempt is time rather than money. That makes the sensible strategy obvious: try the cheap engine first with the photograph you have, look at what comes back, and only reshoot or move up a quality tier once you know which of the six things above was the problem.

Generating four variants at once is the other move worth knowing. When you are exploring rather than producing, four attempts at once tells you more than four sequential ones, because you can see the range rather than one point in it.

Share this X LinkedIn
Read next
Turn an image into a 3D model