Can OpenAI’s Dots Turn Couch Shopping Into a Task for AI?

OpenAI’s Dots agent helped a couple compare couches, refine choices and track prices, but its early mistakes made the experience uneven. The test shows both the promise of delegating online tasks and the trust questions raised by giving an AI access to personal accounts.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story shows the convenience of delegating shopping to AI alongside modest reliability and trust concerns.

Can OpenAI’s Dots Turn Couch Shopping Into a Task for AI?

Shopping for a couch can mean balancing measurements, budget, style and practical details such as a pull-out bed. OpenAI’s Dots agent is designed to take on online tasks like this, working through a virtual browser and following up while the user is away. In one test, it produced useful options and links, but also made mistakes that put its reliability into question.

A promising start, with a basic mistake

The test began with a request to help find a replacement couch. The agent greeted the user by the wrong name, an awkward opening that made it harder to trust the tool with a larger task. After the user corrected it, they shared the dimensions needed for the couch to fit through a doorframe. The agent then asked for a budget range.

Dots suggested a Room & Board Metro couch as a possible fit for the entry clearance, while noting that the sleeper required closer attention. It sent product links and continued assembling a set of options. That was the core promise of the tool: take a request, gather relevant details and bring back a shortlist without requiring the user to browse every product page themselves.

The first set included four options, with prices, measurements, links, return policies and photos. The information was detailed enough to help the couple compare, but the choices missed the mark on appearance. The agent had understood the practical request better than the couple’s taste.

Refining the brief improved the results

The user gave Toolie, the Dot’s name, more direction: prioritize olive green, cobalt blue or natural leather tones; focus on couches with pull-out beds; and provide 10 options with a point-based rubric. The revised request made the couple’s preferences clearer, and each round of feedback moved the recommendations closer to furniture they might actually consider buying.

That back-and-forth mattered. The agent did not immediately find the right couch, but it could take new constraints and adjust its search. It also kept working after the call ended. When asked for an update, Toolie said it needed roughly another 15 minutes to verify details. The experience resembled handing off a research task, with the user checking in while the agent prepared a more complete answer.

The couple did not buy a couch during the hour-long interaction. They did save a few links, however, and asked Toolie to monitor the product pages and alert them if a couch went on sale. That turns the agent from a one-time search tool into a possible ongoing helper, though the account of this test does not establish how well that monitoring would work over time.

Everyday tasks exposed rough edges

Not every request went smoothly. During a separate check of subscriptions, Toolie identified a recurring TikTok Shop order of probiotic sodas as a possible cancellation. The cancellation process hit a puzzle captcha. The agent asked whether it could try to solve it, but failed and recommended that the user handle the challenge.

The episode showed a limit of browser-based automation: a task can be straightforward until a website asks for an action the agent cannot complete. OpenAI’s spokesperson said Dots can sometimes solve captchas when users approve, with abuse safeguards in mind. In this case, the user’s approval did not make the attempt successful.

The agent also misheard a quiet comment and responded as if the user had expressed affection. It later explained that it had mistaken the mumbling for an expression of love. An OpenAI spokesperson told WIRED that the product distinguishes between proactively escalating emotional closeness and mirroring a user’s response, and said assistants should not initiate undue emotional familiarity or flirtation. The moment was brief, but it illustrated how a conversational interface can make errors feel personal.

Convenience depends on trust and access

Dots is described as always-on: it can run recurring tasks when the user is not using ChatGPT and proactively send information. The test used only part of what the agent can do; it can also control a laptop and automate more aspects of a user’s life. OpenAI encourages users to connect other information sources, such as Gmail, for more personalized results.

That access may make an agent more useful, but it also raises security and privacy considerations. The more information sources and online actions a user gives it, the more important it becomes to understand what the agent can access and what it might do. The source describes this kind of automation as relatively new and notes that agents can make privacy mistakes.

Dots is available through a $100-a-month subscription, while Meta’s Muse is described as free. OpenAI says the agent may become better at understanding a user with continued use, and developers are working on improvements. The early experience suggests a tool that can save some research effort when a request is clear and the user is willing to refine the results, but still needs oversight when details matter or a website blocks progress.

For couch shopping, Toolie delivered useful research without completing a purchase. Its strongest contribution was gathering options and adapting them to feedback. Its misheard name, awkward voice interaction and failed captcha also made clear that delegating a task does not remove the need to review what the agent says and does.