Meta tags:
description= As interest grows in world models that predict future states from current observations and actions, accurately modeling part-level dynamics has become increasingly relevant for various applications. Existing approaches, such as Puppet-Master, rely on fine-tuning large-scale pre-trained video diffusion models, which are impractical for real-world use due to the limitations of 2D video representation and slow processing times. To overcome these challenges, we present PartRM, a novel 4D reconstruction framework that simultaneously models appearance, geometry, and part-level motion from multi-view images of a static object. PartRM builds upon large 3D Gaussian reconstruction models, leveraging their extensive knowledge of appearance and geometry in static objects. To address data scarcity in 4D, we introduce the PartDrag-4D dataset, providing multi-view observations of part-level dynamics across over 20,000 states. We enhance the model’s understanding of interaction conditions with a multi-scale drag embedding module that captures dynamics at varying granularities. To prevent catastrophic forgetting during fine-tuning, we implement a two-stage training process that focuses sequentially on motion and appearance learning. Experimental results show that PartRM establishes a new state-of-the-art in part-level motion learning and can be applied in manipulation tasks in robotics.;
keywords= PartRM, 3D reconstruction, Drag;
Headings (most frequently used words):
the, of, to, multi, view, our, and, drags, are, network, stage, in, as, partrm, part, with, state, method, we, first, images, by, designed, drag, module, using, learns, ground, truth, deformed, 3d, supervision, gaussian, results, object, modeling, level, dynamics, large, cross, reconstruction, model, abstract, overview, leverage, fine, tuned, zero123, generate, followed, propagation, distribute, on, moving, parts, then, fed, into, where, embedded, scale, embedding, subsequently, concatenate, unet, down, blocks, adopt, two, training, approach, motion, gaussians, which, stored, database, constructed, lgm, second, appearance, renderings, serving, generation, given, single, image, an, articulated, user, defined, operations, specifying, starting, point, target, endpoint, movable, joints, input, generates, splatting, representation, terminal, after, movement, bibtex, accepted, cvpr, 2025,
Text of the page (most frequently used words):
the (19), and (12), partrm (8), first (8), part (7), multi (7), are (6), that (6), our (6), level (6), hao (6), drag (6), view (6), models (6), page (5), dynamics (5), this (4), with (4), large (4), #reconstruction (4), gao (4), stage (4), #motion (4), appearance (4), for (4), university (4), using (3), which (3), from (3), project (3), you (3), research (3), modeling (3), model (3), 2025 (3), puppet (3), master (3), nvs (3), method (3), object (3), gaussian (3), state (3), results (3), fine (3), images (3), module (3), drags (3), network (3), scale (3), was (2), website (2), author (2), mingju (2), yike (2), pan (2), huan (2), ang (2), zongzheng (2), zhang (2), wenyi (2), dong (2), tang (2), zhao (2), dragapart (2), diffeditor (2), lpips (2), ssim (2), psnr (2), partdrag (2), representation (2), designed (2), embedding (2), two (2), training (2), learns (2), ground (2), truth (2), deformed (2), supervision (2), world (2), future (2), states (2), observations (2), tuning (2), video (2), geometry (2), static (2), data (2), dataset (2), learning (2), code (2), cvpr (2), institute (2), tsinghua (2), computer (2), science (2), indicates (2), built, adopted, free, borrow, just, ask, link, back, footer, licensed, under, creative, commons, attribution, sharealike, international, license, nerfies, academic, template, find, work, useful, your, please, consider, citing, article, gao2025partrm, title, journal, year, bibtex, quantitative, comparison, different, methods, 0758, 9209, 0356, 9531, ours, 245, 361, 0528, 9475, 187, 0579, 9447, 119, 0885, 9004, 0567, 9454, 139, 0873, 8915, 0690, 9343, 128, 0842, 9079, 0918, 9174, 151, 0902, 8988, 1138, 8974, time, objaverse, animation, setting, given, single, image, articulated, user, defined, operations, specifying, starting, point, target, endpoint, movable, joints, input, generates, splatting, terminal, after, movement, generation, leverage, tuned, zero123, generate, followed, propagation, distribute, moving, parts, then, fed, into, where, embedded, subsequently, concatenate, unet, down, blocks, adopt, approach, gaussians, stored, database, constructed, lgm, second, renderings, serving, overview, interest, grows, predict, current, actions, accurately, has, become, increasingly, relevant, various, applications, existing, approaches, such, rely, pre, trained, diffusion, impractical, real, use, due, limitations, slow, processing, times, overcome, these, challenges, present, novel, framework, simultaneously, builds, upon, leveraging, their, extensive, knowledge, objects, address, scarcity, introduce, providing, across, over, 000, enhance, understanding, interaction, conditions, captures, varying, granularities, prevent, catastrophic, forgetting, during, implement, process, focuses, sequentially, experimental, show, establishes, new, art, can, applied, manipulation, tasks, robotics, publicly, available, facilitate, abstract, arxiv, paper, industry, air, department, electrical, engineering, michigan, school, peking, interdisciplinary, information, sciences, iiis, beijing, academy, artificial, intelligence, baal, corresponding, equal, contribution, accepted, cross,
Text of the page (random words):
partrm project page partrm modeling part level dynamics with large cross state reconstruction model accepted to cvpr 2025 mingju gao 1 yike pan 1 2 huan ang gao 1 5 zongzheng zhang 1 wenyi li 1 hao dong 3 hao tang 3 li yi 4 hao zhao 1 5 1 institute for ai industry research air tsinghua university 2 department of electrical engineering and computer science university of michigan 3 school of computer science peking university 4 institute for interdisciplinary information sciences iiis tsinghua university 5 beijing academy of artificial intelligence baal indicates equal contribution indicates corresponding author paper code cvpr 2025 arxiv dataset models abstract as interest grows in world models that predict future states from current observations and actions accurately modeling part level dynamics has become increasingly relevant for various applications existing approaches such as puppet master rely on fine tuning large scale pre trained video diffusion models which are impractical for real world use due to the limitations of 2d video representation and slow processing times to overcome these challenges we present partrm a novel 4d reconstruction framework that simultaneously models appearance geometry and part level motion from multi view images of a static object partrm builds upon large 3d gaussian reconstruction models leveraging their extensive knowledge of appearance and geometry in static objects to address data scarcity in 4d we introduce the partdrag 4d dataset providing multi view observations of part level dynamics across over 20 000 states we enhance the model s understanding of interaction conditions with a multi scale drag embedding module that captures dynamics at varying granularities to prevent catastrophic forgetting during fine tuning we implement a two stage training process that focuses sequentially on motion and appearance learning experimental results show that partrm establishes a new state of the art in part level motion learning and can be applied in manipulation tasks in robotics our code data and models are publicly available to facilitate future research method overview of partrm we first leverage a fine tuned zero123 to generate multi view images followed by our designed drag propagation module to distribute drags on the moving parts the drags and multi view images are then fed into our designed network where the drags are embedded using our multi scale embedding module and subsequently concatenate to the unet down blocks we adopt a two stage training approach in the first stage the network learns part motion using ground truth deformed 3d gaussians as supervision which are stored in the gaussian database constructed by lgm in the second stage the network learns appearance with ground truth deformed multi view renderings serving as supervision results generation results given a single view image of an articulated object and user defined drag operations specifying the starting point and target endpoint of movable joints as input our method generates a 3d gaussian splatting representation of the object s terminal state after movement method setting partdrag 4d objaverse animation hq time psnr ssim lpips psnr ssim lpips diffeditor nvs first 22 52 0 8974 0 1138 19 24 0 8988 0 0902 33 6s 151 2s diffeditor drag first 22 34 0 9174 0 0918 19 46 0 9079 0 0842 11 5s 128 8s dragapart nvs first 24 27 0 9343 0 0690 19 38 0 8915 0 0873 21 4s 139 7s dragapart drag first 24 91 0 9454 0 0567 19 44 0 9004 0 0885 8 5s 119 4s puppet master nvs first 24 20 0 9447 0 0579 64 9s 187 5s puppet master drag first 24 42 0 9475 0 0528 245 8s 361 5s partrm ours 28 15 0 9531 0 0356 21 38 0 9209 0 0758 4 2s quantitative comparison of different methods bibtex if you find our work useful in your research please consider citing article gao2025partrm title partrm modeling part level dynamics with large 4d reconstruction model author mingju gao yike pan huan ang gao zongzheng zhang wenyi li hao dong hao tang li yi hao zhao journal year 2025 this page was built using the academic project page template which was adopted from the nerfies project page you are free to borrow the of this website we just ask that you link back to this page in the footer this website is licensed under a creative commons attribution sharealike 4 0 international license
|