Skip to content

Image-to-3D: MV-Adapter (Apache-2.0) as a multi-view texture upgrade for TripoSG #805

Description

@fernandotonon

Image-to-3D: MV-Adapter (Apache-2.0) as a multi-view texture upgrade for TripoSG

Motivation

TripoSG (#764) gives us the best local single-image geometry, but it's geometry-only. Today its colour comes from (a) projecting the input photo onto the visible front and (b) inferring the back from TripoSR's colour field. The front is photo-accurate; the back is approximate and can be blotchy where the two models' scales disagree.

MV-Adapter (huanngzh/MV-Adapter, "Multi-view Consistent Image Generation Made Easy") generates a set of consistent multi-view images conditioned on geometry — the same technique Meshy/Tripo use before baking. Feeding those views into our existing MultiViewTextureBaker would give a genuinely good, consistent texture on all sides, not just the front.

This is the tracked follow-up from the #764 texture work (the generate3d "Generate texture (AI)" option currently uses a photo-front + SD-txt2img-back multi-view bake; MV-Adapter replaces the back-generation with proper multi-view-consistent generation).

License due-diligence (vetted — GO)

  • Code: Apache-2.0 (GitHub LICENSE, verbatim Apache 2.0).
  • Adapter weights: Apache-2.0 (HF huanngzh/mv-adapter model card). Variants: t2mv/i2mv/tg2mv/ig2mv for SD2.1 and SDXL bases; the geometry-guided tg2mv/ig2mv SDXL checkpoints are the texturing ones.
  • Training data: ~70K filtered Objaverse subset (ODC-BY-1.0) + Cap3D captions (permissive). No non-commercial dataset; no UltraSharp/ambientCG-style provenance discrepancy.
  • Base model (the caveat): MV-Adapter is an adapter — it needs an SD2.1 or SDXL base, both CreativeML Open RAIL++-M. This is commercial-and-redistribution-OK with ethical use-restrictions, and is the same RAIL license family we already ship via the SD1.5 stable-diffusion.cpp path — no new licensing category. NOT in any rejected bucket (SF3D/Zero123 NC, Hunyuan3D EU-excluded, GPL deps).
  • Action: if pursued, add MV-Adapter + the chosen base to THIRD_PARTY_AI_MODELS.md and propagate the RAIL Attachment-A notice as we do for SD1.5.
  • ⚠ No SD1.5 variant — smallest base is SD2.1 (512). Reusing the existing SD1.5 sd.cpp path is not possible without adding SD2.1/SDXL weights.

Output shape (fits our baker)

MV-Adapter emits 6 multi-view images (768²), NOT a textured mesh. Our MultiViewTextureBaker already projects per-view images onto a mesh's UV0 with facing-weight blending + colour-match + dilation — so MV-Adapter's views drop straight in. The ig2mv (image+geometry→multi-view) variant is the right one: condition on the input photo and the TripoSG geometry (via rendered depth/normal from MeshDepthRenderer), get back-consistent views, bake.

Implementation options (pick during the issue)

Option A — ONNX export (mirrors TripoSG/TripoSR/UniRig)

Export the adapted UNet step graph + VAE + text/image encoders, run an SD-style scheduler loop in C++ (we already do this for TripoSG's flow loop).

  • Cost: HIGH (~2–4 weeks). SDXL-scale UNet (~2.6B) needs int8 tiers; the decoupled multi-view attention is a non-standard module the exporter must trace; texture pipeline extras (LaMa inpaint — new; RealESRGAN — we have via AI: Real-ESRGAN texture upscaling #405). No upstream ONNX exists.
  • Consistent with the project's "no torch at runtime" stance.

Option B — teach stable-diffusion.cpp to load the adapter

Load MV-Adapter's attention layers onto an SD2.1/SDXL base inside sd.cpp.

  • sd.cpp has no MV-Adapter support today (also non-trivial C++), but avoids a bespoke ONNX export and reuses the shipped SD runtime. Would need SDXL support in our sd.cpp integration.

Proposed slices

  1. Spike / export feasibility — export the ig2mv SDXL adapted-UNet to ONNX (or prototype the sd.cpp adapter-load); measure size, verify a single denoise step round-trips. Decide A vs B.
  2. Runtime — C++ multi-view scheduler loop producing the 6 views from (photo + geometry depth/normal). int8 tier.
  3. Bake wiring — feed views → MultiViewTextureBaker (already exists); replace the txt2img-back path in the generate3d "Generate texture (AI)" flow. Keep the current photo-front + SD-back as the fallback when MV-Adapter models aren't present.
  4. Hosting + docs — host the exported graphs on the HF models repo (download-on-first-use, QTMESH_MVADAPTER_* env/QSettings overrides, offline guard); THIRD_PARTY_AI_MODELS.md.

Acceptance

  • Generating a texture for a TripoSG mesh produces a consistent texture on all sides (no back blotching), noticeably better than the current photo-front + txt2img-back.
  • Redistribution-clean (Apache-2.0 adapter + RAIL base, documented).
  • Fails soft to the existing multi-view bake when MV-Adapter models are unavailable.
  • ENABLE_STABLE_DIFFUSION-guarded; no runtime torch.

Notes

Activity

  1. added 5 commits that reference this issue on Jul 8, 2026
  2. added a commit that references this issue on Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions