Reconstructing a person in motion as a detailed 3D figure requires more than a collection of still images. HumanRF is a method designed to learn from moving human footage and produce high-resolution reconstructions that remain consistent over time. Its creators are also releasing ActorsHQ, a dataset built to capture that motion in detail.
A dataset built for moving people
Neural Radiance Fields, or NeRFs, learn 3D representations from photos or video. They can be used to render objects and scenes, and some approaches focus on scenes or objects that move. HumanRF applies this kind of representation to people in motion.
The method is trained on ActorsHQ, a dataset containing 39,765 frames of dynamic human motion. The footage was captured using multiple cameras, alongside an LED array for global illumination. This setup was designed to collect higher-resolution material than older datasets described in the report, which reach a maximum resolution of 4 MP.
ActorsHQ includes four females and four males performing 20 randomly selected motions. The capture system uses 160 Ximea 12 MP cameras operating at 25 frames per second, with an illumination array of 420 LEDs. Together, the footage and equipment provide the visual input for learning how a person’s appearance changes as they move.
Keeping detail consistent over time
HumanRF aims to capture the high-resolution data in ActorsHQ while reconstructing actors consistently across sequences. That consistency matters because an output that looks convincing in one frame may still fail to represent a continuous movement well. The method is intended to preserve both fine visual details and changes in pose over time.
The team drew inspiration from Nvidia's Instant-NGP, adding a time dimension to the encodings used in that approach. In practical terms, the added dimension helps the representation account for motion as it learns from a sequence, rather than treating each view as an isolated image.
This focus distinguishes HumanRF from a reconstruction aimed only at a static portrait. The goal is to represent a human actor through a longer span of movement, with the visual detail retained as the body changes position. The article describes the resulting reconstructions as high quality and temporally consistent, including over long sequences.
What 3D avatars could make possible
NeRFs are viewed as a technology with potential applications in 3D graphics, video conferencing and, in the future, the metaverse. HumanRF and ActorsHQ could help advance photorealistic reconstruction of virtual humans by giving researchers both a method and a dataset to work with.
There is also a possible product direction for Synthesia, the synthetic media AI startup involved in the work. The team says it plans to explore ways to control the articulation of trained actors. If that work succeeds, it could help move the company's products from 2D recordings toward dynamic 3D avatars.
That possibility depends on further research. HumanRF currently focuses on learning reconstructions from captured motion; controlling how a trained actor articulates is a planned area of exploration, rather than a capability the article says is already available.
Research resources and next steps
The team says it has released the ActorsHQ dataset and plans to make the code and dataset available on the HumanRF project website. Those resources could let others examine the capture data and explore how the reconstruction method performs.
For now, the work brings together high-resolution multi-view footage and a NeRF method adapted to motion. Its significance lies in the combination: detailed recordings supply the training material, while the time-aware representation aims to preserve the actor’s appearance across movement. Further work on controllable articulation may show how far this approach can extend toward usable 3D avatars.