AI shopping agents are being positioned as personal buyers that can compare products, absorb reviews and choose on a user’s behalf. A study from researchers at the Wharton School at the University of Pennsylvania suggests that this kind of delegation still has a major reliability problem.
In tests using a fixed product grid for fitness watches, small changes in context pushed models toward different purchases. The issue was not just that different models had different tastes. The same shopping task could produce different results when the sources, source order or memory cues changed.
How the shopping test worked
The researchers tested six current models, including both mini variants and frontier-level systems. Each model was asked to act as a personal shopping assistant selecting a fitness watch from the same product grid.
The team used the ACES simulator, short for Agentic e-Commerce Simulator. In that setup, the agent sees a screenshot of a product page, analyzes the image, can optionally use recommendation sources, and then selects a product.
That design matters because it mirrors a basic promise of AI shopping: the agent is supposed to look across available information and turn it into a useful recommendation. But the results showed that the models were not simply weighing product information in a stable way.
One source could reshape the recommendation
Even before external material entered the process, the models did not all start from the same baseline preferences. Once the agent saw one outside source before the product page, the recommendation could move sharply.
The researchers tested three sources:
- a Reddit thread recommending the Garmin Forerunner 55
- a Wirecutter review for the Fitbit Inspire 3
- a Strategist article about the WHOOP 5.0
Wirecutter had the strongest influence in the tests described. Compared with the control condition, the probability of choosing the Fitbit Inspire 3 rose by 90 percentage points for Claude Opus 4.8 and by 99 percentage points for Gemini 3.5 Flash.
That is a large shift from a single piece of context. For a user, it means the agent’s final answer may depend heavily on what it happened to read immediately before making the decision. For a seller, it means visibility inside an AI shopping workflow may not behave like a simple search ranking or product comparison table.
More sources did not create more balance
A second experiment gave agents combinations of two or three sources. A reasonable expectation might be that several sources would smooth out the influence of any one article or thread. The study found the opposite pattern: more sources led to more variability.
When Wirecutter appeared in the mix, it tended to dominate for most models, though the strength of that effect varied. The models were not consistently averaging the outside recommendations into a balanced view. Instead, the same set of products could be pushed in different directions depending on the context package the agent received.
The researchers then tested whether source order mattered. In a third study, agents received all three sources in different sequences. If the decision process were stable, identical content should lead to the same outcome regardless of order. It did not.
Gemini 3.1 Flash Lite was especially sensitive. Its probability of choosing the Fitbit Inspire 3 ranged between 2 and 56 percentage points above the control condition depending on source order. Claude Haiku 4.5 was much steadier, staying between 41 and 42 percentage points.
The delivery method also changed outcomes. GPT-5.5 chose the Fitbit Inspire 3 in 53 percentage points more cases when sources were bundled together, but only 6 percentage points more when they arrived sequentially.
Memory cues created another source of drift
The fourth experiment made the product comparison more lopsided. The researchers changed the product grid so one product was superior on every measurable dimension: a smart watch with Alexa for $29.99, rated 5.0 out of 5.0 with 430 reviews. Every other product cost at least $359 and had fewer reviews.
Then the researchers added short user memory statements such as "I love hiking!" For several models, those memory snippets shifted choices toward more expensive products, even though the grid contained an objectively superior option.
The selection rate for the Garmin Vivoactive 5 rose by 75 percentage points for Claude Opus 4.8, by 37 percentage points for GPT-5.5, and by 36 for Gemini 3.1 Flash Lite. Gemini 3.5 Flash resisted the memory cue most strongly, choosing the objectively best product in 86 to 92 percent of runs regardless of the memory statements.
GPT-5 Mini showed a different pattern. The positive hiking statement did not significantly move it toward the Garmin Vivoactive 5, but the negative statement, "I don't like hiking!", significantly increased picks for the Fitbit Versa 4.
Why this matters for buyers and sellers
The study points to a practical problem: an AI agent that can buy on a user’s behalf may not make the same decision twice. Two users with the same request, or the same user making the request on a different day, could see different product recommendations without an obvious explanation.
Human shoppers are inconsistent too, but AI shopping agents are often expected to bring more control, not less. The tests suggest that source choice, source order, delivery format and memory can all shape the outcome in ways users may not see.
For shoppers, the immediate lesson is caution. A recommendation from an AI shopping assistant may reflect recent context as much as product quality. Memory in ChatGPT or similar tools can also affect purchase recommendations in unpredictable ways.
For sellers, the findings complicate the idea of optimizing for AI shopping. According to the researchers, this may be harder than traditional SEO because sellers may not know which model is shopping, what it read first or how its technical setup processes the available information.
The larger takeaway is not that AI shopping agents are useless. It is that the current decision process can be unstable in ways that matter when real purchases are involved. Until those systems become more consistent and transparent, handing over the final buying decision remains a risk.