A reflection in the eye can carry clues about the space around a person. Researchers at the University of Maryland, College Park, explored how AI could use those clues to rebuild a 3D scene, even when a single camera is pointed at the person rather than directly at the scene.
Reading the reflection through the cornea
The project, "Seeing the World through Your Eyes," treats eye reflections as a source of visual information about the surrounding world. The researchers describe this as an "underappreciated source of information about what the world around us looks like". Their approach uses a NeRF-based method to infer a scene from reflections captured in the eye.
That task requires more than spotting a reflection. The system has to estimate where the eye is and which way it is oriented, while telling reflected details apart from the iris's own patterns. The team uses the geometry of the cornea, which is fairly uniform in healthy adults, to estimate the eye's position and orientation from its size in an image.
Finding the cornea accurately also matters. The researchers refine an initial position estimate by optimizing cornea detection, a step they say is critical to making the method robust. Errors in locating the cornea or estimating its shape can affect how the system interprets the reflection.
Separating scenery from iris texture
The iris and the reflected scene appear together in the image, so the system must learn which visual patterns belong to which source. To address this, the researchers adapted a training framework from Nerfstudio. They trained NeRF to represent both the reflected 3D scene and the iris pattern, while learning to separate the two.
The method relies on a distinction between how these elements behave across viewpoints. It assumes the iris has a general radially symmetric shape and that its texture stays the same as the viewing perspective changes. The reflection, by contrast, changes as the perspective shifts. Those assumptions give the model a way to assign observed details to either the eye or the surrounding scene.
This separation is central to the reconstruction. If the model mistakes iris texture for part of the scene, or the other way around, the resulting 3D representation may be less accurate. The researchers identify this as a complicated part of the process rather than a simple image-reading step.
One camera, changing viewpoints
NeRF reconstructions typically use multiple camera views of the object or space being reconstructed. In this work, the camera observed a person moving through its field of view. Small movements changed what was reflected in the eyes, providing different perspectives on the environment even though the scene itself was not viewed directly by multiple cameras.
The team evaluated the approach with synthetic eye images rendered with Blender and with real photographs of a person moving within the camera's field of view. In tests using synthetic eye models, it was able to produce full-scene reconstructions based only on eye reflections. The real-image experiments also showed promise, despite the low resolution and challenges in estimating cornea position and geometry.
The setup highlights how motion can create useful visual variation. As a person moves, the reflection shifts, giving the system information from more than one perspective. In principle, that changing reflection can supply some of the viewpoints that a conventional multi-camera reconstruction would seek.
What the early results do and do not show
The work was tested in a laboratory, and the researchers caution that real-world conditions introduce other factors. The assumptions about iris texture may be too simple for some situations. Brighter iris textures or strong eye rotation could make it harder to separate the reflection from the eye itself.
Image quality and geometric estimates also limit the method. The researchers note inaccuracies in locating the cornea and estimating its geometry, as well as the inherently low resolution of the images. These constraints matter because the reflection is a small visual signal, and errors in interpreting the eye can change how the reconstructed scene is formed.
The results therefore demonstrate a possible technique, not a dependable way to recover a scene from every photograph of a person. The researchers hope the work encourages further study of how unexpected visual signals can reveal information about the world around us. Eye reflections are one example of a signal that might be useful when viewed from a new computational perspective.