GPT-4 Gives Spot a Voice, but Not Real-Time Answers

Boston Dynamics equipped Spot with GPT-4, speech recognition, a speaker, microphone and camera arm to hold conversations and act as a tour guide. The experiments produced surprising responses, but also errors and a roughly six-second delay, underscoring that fluent language does not mean human understanding.

WTF Index NEUTRAL
◄ Terminator 2 Idiocracy 2 ►

The demo gives Spot a more capable conversational interface, but errors and latency underline its limits without showing clear harm or human skill erosion.

GPT-4 Gives Spot a Voice, but Not Real-Time Answers

Spot, Boston Dynamics’ robot dog, can now answer questions during tours of the company’s headquarters. A demonstration pairing the robot with GPT-4 shows how language models can give a physical machine a conversational interface—and how easily that can appear more capable than it is.

How Spot holds a conversation

Boston Dynamics added a Bluetooth speaker and microphone to Spot, along with a camera-equipped arm that serves as a head and neck. The robot’s grasping hand opens and closes like a talking mouth, giving its spoken replies a visible gesture.

For language and image processing, the system uses OpenAI’s GPT-4 model, Visual Question Answering datasets and OpenAI’s Whisper speech recognition software. A person can begin an interaction with “Hey, Spot!” and ask questions as the robot performs its tour guide duties.

Spot can also recognize a person nearby and turn toward them during a conversation. That combination of speech, camera input, movement and a speaking gesture makes the exchange feel more embodied than a voice assistant answering from a stationary device.

Unexpected answers reveal the model’s associations

Boston Dynamics reported responses it described as emergent capabilities: behaviors that were not explicitly programmed as instructions. Asked about Boston Dynamics founder Marc Raibert, Spot said it did not know him and suggested seeking support from the IT help desk. The prompt had not specifically told Spot to make that suggestion.

Asked about its parents, Spot referred to “Spot V1” and “Big Dog,” earlier robots in the lab, as its elders. The robot also maintained a defined character and regularly made snarky remarks, according to Boston Dynamics.

These moments can sound like reasoning, but the article cautions against treating them as evidence of consciousness or humanlike intelligence. The examples instead show how a language model can connect ideas such as asking a question with a help desk, or parents with age, and produce a response that fits the conversation.

Fluent speech still comes with limits

The demonstration also had clear weaknesses. Spot sometimes gave incorrect answers, including describing the logistics robot “Stretch” as being for yoga. And a reply took about six seconds to arrive, making the interaction far from real time.

That delay matters for how people experience a talking robot. A conversational pause can make an exchange feel less immediate, while a confident but mistaken answer can make it harder to judge what the system actually knows. The polished delivery and body language may make an answer seem more deliberate than it is.

Boston Dynamics based its implementation on Microsoft's ChatGPT rules for robots. The company blog includes a detailed description of the prompts. Together, those choices shape the character and responses that people encounter, while the reported errors show that a carefully framed persona does not guarantee accuracy.

Why put a language model in a robot?

Boston Dynamics did not describe specific plans for the modified Spot. It did, however, point to possible uses for language-enabled robots, including tourist guides, customer service and personal companions. If people can teach or direct a robot through conversation, that could also shorten the learning curve for new tasks.

The experiment illustrates a broader idea: language models can connect their store of world knowledge to robots operating in physical settings. The source points to Google’s RT-2 and SayCan projects as other examples of this direction. In principle, language can help a robot respond to requests and carry out complex actions without each one needing additional programming.

Spot’s tour guide performance offers a small-scale look at that possibility. It also makes the distinction between conversational fluency and dependable capability visible: the robot can speak, gesture and improvise, but its answers can still be wrong and slow.