Voice commands turn into movable objects in a ChatGPT AR app

Developer Bart Trzynadlowski used ChatGPT to generate JavaScript for an augmented reality app called ChatARKit. Users can ask it to find and place Sketchfab models, though identical requests can produce different or faulty code.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 2 ►

The app makes spatial tasks easier through voice commands, but unreliable generated code mildly raises concerns about dependence and quality.

Voice commands turn into movable objects in a ChatGPT AR app

A spoken request can become an object on a desk or floor in ChatARKit, an augmented reality app built by developer Bart Trzynadlowski. The app uses ChatGPT to generate code that selects and manipulates 3D models in response to natural language commands.

From a spoken request to an AR object

ChatARKit connects voice input with a JavaScript environment. The app uses OpenAI's Whipser to recognize what the user says, then passes the command to ChatGPT as a prompt. ChatGPT generates code intended to carry out the request.

That code can search Sketchfab for a matching 3D object and place it on a desktop or floor. Users can also ask for changes such as scaling or rotating a model. The interaction is framed as a direct instruction: describe the object and what it should do, and the system attempts to translate that into an action in the scene.

Examples reported by Trzynadlowski include placing a cube on the nearest plane, setting a spinning cube on the floor, and putting a sports car on a table before rotating it 90 degrees. A more elaborate request asks for a school bus on the nearest plane and has it drive back and forth along the surface.

Useful demonstrations, uneven results

The examples show how language can serve as an interface for spatial computing. A user does not need to select a model from a list and then separately specify every adjustment; a single instruction can combine an object, a location and an action.

But the system is not dependable, according to Trzynadlowski. The same command can lead ChatGPT to produce substantially different output on separate attempts. Sometimes it inserts incorrect JavaScript lines into the app, interrupting the action the user asked for.

There is also a mismatch risk between the words in a request and the way an object is looked up. Trzynadlowski says ChatGPT can turn an object's description into a code identifier. When that happens, the app may no longer be able to retrieve the intended 3D model from Sketchfab.

Those problems matter because the app relies on generated code to connect an open-ended instruction to concrete operations: finding a model, positioning it, and changing its properties. A fluent request does not guarantee that each step will work correctly.

Text-to-3D points toward another route

ChatARKit uses existing models from Sketchfab. Another approach described in the source is to generate 3D representations from text. Developer Jasmine Roberts demonstrated a VR implementation using OpenAI's Point-E, which produces 3D point clouds representing models from text input.

The source describes Point-E as a starting point for OpenAI's work on text-to-3D synthesis. Roberts' demo ran in real time, while Point-E generation was described as taking about one to two minutes on a single Nvidia V100 GPU. The two examples illustrate different parts of a possible workflow: one turns language into instructions for placing existing objects; the other creates a 3D representation from a description.

Other text-to-3D systems mentioned in the source include Google's Dreamfusion and Nvidia's Magic3D. Together, these projects point to ways natural language could make 3D content easier to create or use. The source connects wider availability of 3D content to the metaverse thesis, while presenting that possibility as a future prospect.

An experiment in controlling 3D scenes with language

ChatARKit is available for free as open source on Github, according to the source. Its sample commands make the idea easy to grasp: tell the app what to place, where to put it, and how to move or transform it.

The current limitations are just as clear. ChatGPT may vary its response, generate faulty JavaScript, or interfere with retrieving the requested model. The demonstration therefore shows both the appeal of natural language control for AR and the reliability challenge that comes with asking a language model to produce executable code.