Video can show how an object looks from different angles, but reconstructing its shape does not explain how it moves or what physical properties govern that motion. PAC-NeRF is a research approach designed to learn both from video, combining a 3D scene representation with physical simulation.
Learning appearance and physical behavior together
Neural radiance fields, or NeRFs, learn to represent and render the geometry and lighting of scenes from video. They are used for tasks such as video production and 3D reconstruction, but conventional NeRF scenes are static and do not capture objects' physical properties.
PAC-NeRF, short for Physics Augmented Continuum Neural Radiance Fields, aims to address that gap. Researchers from UC Los Angeles, University of Maryland, MIT CSAIL, Columbia University, UMass Amherst, and the MIT-IBM Watson AI Lab developed the approach to learn an object's geometric structure alongside its physical properties.
That combination allows the representation to describe dynamic scenes. The learned properties may also be useful in other pipelines, though the source does not specify those future uses.
Physics constrains what the model can learn
The PAC-NeRF architecture is designed to follow conservation laws from continuum mechanics during training. These constraints are intended to keep the model's inferred states physically plausible, rather than allowing it to explain video with motion that violates the modeled physics.
The method combines neural rendering with the material point method, or MPM, for differentiable physics simulation. In practical terms, this links the scene's rendered appearance to a simulation of how its material moves and changes.
PAC-NeRF can represent objects with properties associated with Newtonian and non-Newtonian fluids, as well as elastic materials, plasticine, and sand. This range suggests the method is intended for more than rigid objects, where motion and deformation are central to what a reconstruction needs to capture.
Training and capture still have constraints
The reported processing times show that building a PAC-NeRF representation takes substantially longer than producing an individual view. Training a mesh currently takes 1.5 hours on an Nvidia 3090 GPU, while rendering one image takes about a second.
The capture setup also limits what the current method can handle. It requires fixed camera angles and synchronized, calibrated cameras. The scene should also have a simple background that can be removed easily.
Those requirements help explain why many of the examples use footage from physical simulations. The researchers also show a real-world example, but the stated camera and background conditions remain important considerations for applying the approach.
More materials are a future direction
The research team describes PAC-NeRF as an improvement over earlier approaches that either do not learn geometric structure or do not incorporate physical laws. That comparison reflects the method's goal of bringing geometry, appearance, and physical behavior into a shared representation.
Future work is planned to extend the MPM framework to objects with other physical properties. The examples named by the researchers include cloth, stiff materials, and joint body simulations.
For now, PAC-NeRF offers a way to connect what a video shows with a model of how an object behaves. Its usefulness depends on the capture conditions and material types it can represent, while the planned extensions point toward a broader range of dynamic scenes.