Hook
Every AI 3D model you have ever generated came out as one fused blob. The wheels are welded to the car. The drawer is glued shut. KaiNinja just fixed that — it generates the object and its separate parts in a single pass, from one image, with no mask and no segmenter.
The Story
Part-level generation is the quiet bottleneck of AI 3D. Meshy, Tripo, Rodin, TRELLIS — they all hand you a beautiful watertight shell. Then you open it in Blender and discover it is one connected surface. You cannot move the arm. You cannot open the lid. You cannot swap the wheel. For anyone building games, product viz, or anything that needs to move, that is a wall.
KaiNinja, released September 14 by Alaya Lab with the University of Tokyo, UC Merced and Institute of Science Tokyo, attacks that wall head-on. It is built on top of Microsoft’s TRELLIS.2 — the image-to-3D model the Lab covered back in August — and keeps its speed and quality. But it adds one clever idea: a dual-volume O-Voxel.
Here is the problem it solves. TRELLIS.2 stores one sheet of surface per voxel — one tiny cube of space. That works great for the outer skin of an object. But where two parts touch, there are two surfaces crammed into the same spot: the outside of the drawer and the inside of the cabinet. A single volume cannot hold both. So KaiNinja uses two complementary volumes at once. One holds the “even” parts, the other the “odd” ones, and together they can describe the exact seam where parts meet. No 2D mask. No part segmenter. No slow per-object optimization.
The other headline is where the training data came from. Clean, part-labelled 3D assets are rare and expensive. So the team built a pipeline where an LLM-driven agent authors 3D assets part by part — think of an AI writing the CAD script for a chair, leg by leg. To their knowledge, KaiNinja is the first 3D generative model trained on agent-authored part data. That is a genuinely new idea: AI generating the training set for the next AI.
Why You Should Care
The numbers are not subtle. Against other part-generation pipelines, KaiNinja cuts whole-object Chamfer distance — a measure of how far the shape drifts from the truth — by 40%, while raising strict part F-score by 16%. In plain terms: the whole object is more accurate and the parts are cleaner at the same time. Usually you trade one for the other.
- Game devs: assets that arrive pre-separated are assets you can rig, animate, and swap without hours of manual cutting.
- Product & archviz: a cabinet whose drawers actually open, a lamp whose shade detaches — straight from one reference photo.
- 3D artists: parts mean per-part materials, per-part edits, and a real starting point in Blender instead of a sculpting cleanup job.
Try It / Follow Them
Honest note: the code is marked coming soon, so you cannot run KaiNinja tonight. But the project page has 14 interactive 3D results you can rotate and pull apart in your browser, plus a 55-second video showing generated parts with materials and hand-authored animation — worth five minutes of your time.
- Project page & interactive demos: alaya-lab.github.io/KaiNinja
- Paper: arXiv 2609.15659
- Foundation it builds on: our TRELLIS.2 write-up
IK3D Lab Take
For two years AI 3D chased prettier surfaces. KaiNinja is a reminder that the real prize is structure — an object you can take apart is an object you can actually use. The dual-volume trick is elegant, and the agent-authored data idea hints at how these models keep improving once human-labelled 3D runs dry. Once the code lands and someone wires it into a ComfyUI node, this stops being a paper and starts being part of the pipeline. We will be first in line to test it. Watch this one.



