You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Image-to-3D: MV-Adapter (Apache-2.0) as a multi-view texture upgrade for TripoSG #805
Image-to-3D: MV-Adapter (Apache-2.0) as a multi-view texture upgrade for TripoSG
Motivation
TripoSG (#764) gives us the best local single-image geometry, but it's geometry-only. Today its colour comes from (a) projecting the input photo onto the visible front and (b) inferring the back from TripoSR's colour field. The front is photo-accurate; the back is approximate and can be blotchy where the two models' scales disagree.
MV-Adapter (huanngzh/MV-Adapter, "Multi-view Consistent Image Generation Made Easy") generates a set of consistent multi-view images conditioned on geometry — the same technique Meshy/Tripo use before baking. Feeding those views into our existing MultiViewTextureBaker would give a genuinely good, consistent texture on all sides, not just the front.
This is the tracked follow-up from the #764 texture work (the generate3d "Generate texture (AI)" option currently uses a photo-front + SD-txt2img-back multi-view bake; MV-Adapter replaces the back-generation with proper multi-view-consistent generation).
Adapter weights: Apache-2.0 (HF huanngzh/mv-adapter model card). Variants: t2mv/i2mv/tg2mv/ig2mv for SD2.1 and SDXL bases; the geometry-guided tg2mv/ig2mv SDXL checkpoints are the texturing ones.
Training data: ~70K filtered Objaverse subset (ODC-BY-1.0) + Cap3D captions (permissive). No non-commercial dataset; no UltraSharp/ambientCG-style provenance discrepancy.
Base model (the caveat): MV-Adapter is an adapter — it needs an SD2.1 or SDXL base, both CreativeML Open RAIL++-M. This is commercial-and-redistribution-OK with ethical use-restrictions, and is the same RAIL license family we already ship via the SD1.5 stable-diffusion.cpp path — no new licensing category. NOT in any rejected bucket (SF3D/Zero123 NC, Hunyuan3D EU-excluded, GPL deps).
Action: if pursued, add MV-Adapter + the chosen base to THIRD_PARTY_AI_MODELS.md and propagate the RAIL Attachment-A notice as we do for SD1.5.
⚠ No SD1.5 variant — smallest base is SD2.1 (512). Reusing the existing SD1.5 sd.cpp path is not possible without adding SD2.1/SDXL weights.
Output shape (fits our baker)
MV-Adapter emits 6 multi-view images (768²), NOT a textured mesh. Our MultiViewTextureBaker already projects per-view images onto a mesh's UV0 with facing-weight blending + colour-match + dilation — so MV-Adapter's views drop straight in. The ig2mv (image+geometry→multi-view) variant is the right one: condition on the input photo and the TripoSG geometry (via rendered depth/normal from MeshDepthRenderer), get back-consistent views, bake.
Implementation options (pick during the issue)
Option A — ONNX export (mirrors TripoSG/TripoSR/UniRig)
Export the adapted UNet step graph + VAE + text/image encoders, run an SD-style scheduler loop in C++ (we already do this for TripoSG's flow loop).
Cost: HIGH (~2–4 weeks). SDXL-scale UNet (~2.6B) needs int8 tiers; the decoupled multi-view attention is a non-standard module the exporter must trace; texture pipeline extras (LaMa inpaint — new; RealESRGAN — we have via AI: Real-ESRGAN texture upscaling #405). No upstream ONNX exists.
Consistent with the project's "no torch at runtime" stance.
Option B — teach stable-diffusion.cpp to load the adapter
Load MV-Adapter's attention layers onto an SD2.1/SDXL base inside sd.cpp.
sd.cpp has no MV-Adapter support today (also non-trivial C++), but avoids a bespoke ONNX export and reuses the shipped SD runtime. Would need SDXL support in our sd.cpp integration.
Proposed slices
Spike / export feasibility — export the ig2mv SDXL adapted-UNet to ONNX (or prototype the sd.cpp adapter-load); measure size, verify a single denoise step round-trips. Decide A vs B.
Runtime — C++ multi-view scheduler loop producing the 6 views from (photo + geometry depth/normal). int8 tier.
Bake wiring — feed views → MultiViewTextureBaker (already exists); replace the txt2img-back path in the generate3d "Generate texture (AI)" flow. Keep the current photo-front + SD-back as the fallback when MV-Adapter models aren't present.
Hosting + docs — host the exported graphs on the HF models repo (download-on-first-use, QTMESH_MVADAPTER_* env/QSettings overrides, offline guard); THIRD_PARTY_AI_MODELS.md.
Acceptance
Generating a texture for a TripoSG mesh produces a consistent texture on all sides (no back blotching), noticeably better than the current photo-front + txt2img-back.
Image-to-3D: MV-Adapter (Apache-2.0) as a multi-view texture upgrade for TripoSG
Motivation
TripoSG (#764) gives us the best local single-image geometry, but it's geometry-only. Today its colour comes from (a) projecting the input photo onto the visible front and (b) inferring the back from TripoSR's colour field. The front is photo-accurate; the back is approximate and can be blotchy where the two models' scales disagree.
MV-Adapter (huanngzh/MV-Adapter, "Multi-view Consistent Image Generation Made Easy") generates a set of consistent multi-view images conditioned on geometry — the same technique Meshy/Tripo use before baking. Feeding those views into our existing
MultiViewTextureBakerwould give a genuinely good, consistent texture on all sides, not just the front.This is the tracked follow-up from the #764 texture work (the
generate3d"Generate texture (AI)" option currently uses a photo-front + SD-txt2img-back multi-view bake; MV-Adapter replaces the back-generation with proper multi-view-consistent generation).License due-diligence (vetted — GO)
LICENSE, verbatim Apache 2.0).huanngzh/mv-adaptermodel card). Variants:t2mv/i2mv/tg2mv/ig2mvfor SD2.1 and SDXL bases; the geometry-guidedtg2mv/ig2mvSDXL checkpoints are the texturing ones.stable-diffusion.cpppath — no new licensing category. NOT in any rejected bucket (SF3D/Zero123 NC, Hunyuan3D EU-excluded, GPL deps).THIRD_PARTY_AI_MODELS.mdand propagate the RAIL Attachment-A notice as we do for SD1.5.Output shape (fits our baker)
MV-Adapter emits 6 multi-view images (768²), NOT a textured mesh. Our
MultiViewTextureBakeralready projects per-view images onto a mesh's UV0 with facing-weight blending + colour-match + dilation — so MV-Adapter's views drop straight in. Theig2mv(image+geometry→multi-view) variant is the right one: condition on the input photo and the TripoSG geometry (via rendered depth/normal fromMeshDepthRenderer), get back-consistent views, bake.Implementation options (pick during the issue)
Option A — ONNX export (mirrors TripoSG/TripoSR/UniRig)
Export the adapted UNet step graph + VAE + text/image encoders, run an SD-style scheduler loop in C++ (we already do this for TripoSG's flow loop).
Option B — teach
stable-diffusion.cppto load the adapterLoad MV-Adapter's attention layers onto an SD2.1/SDXL base inside sd.cpp.
Proposed slices
ig2mvSDXL adapted-UNet to ONNX (or prototype the sd.cpp adapter-load); measure size, verify a single denoise step round-trips. Decide A vs B.MultiViewTextureBaker(already exists); replace the txt2img-back path in thegenerate3d"Generate texture (AI)" flow. Keep the current photo-front + SD-back as the fallback when MV-Adapter models aren't present.QTMESH_MVADAPTER_*env/QSettings overrides, offline guard);THIRD_PARTY_AI_MODELS.md.Acceptance
ENABLE_STABLE_DIFFUSION-guarded; no runtime torch.Notes
MultiViewTextureBaker(both infeat/triposg-backend/ PR Image→3D: TripoSG backend — 1.5B rectified-flow geometry (MIT), selectable next to TripoSR #794).