Nvidia is adapting its Avatar Cloud Engine (ACE) for video games, where it could let players speak with non-player characters and receive spoken, animated responses. The system combines language, speech, and facial animation tools, with low-latency interaction presented as a key goal.
Three tools form the conversation loop
ACE brings together NeMo, Riva, and Omniverse Audio2Face. NeMo supplies large language models that developers can customize with a character’s backstory and dialogue data, giving the model material to draw on when responding.
Riva handles speech recognition and text-to-audio conversion, supporting live exchanges with NeMo. Audio2Face then turns the resulting audio into facial animation. Nvidia says those animations can be used with MetaHuman characters in Unreal Engine 5 through Omniverse links.
Together, the components connect a player's spoken question with a character's answer and visible expression. That pipeline is intended to make an NPC feel more conversational than a character limited to a fixed set of prerecorded lines.
A demo shows the intended interaction
Nvidia and AI startup Convai demonstrated an ACE-powered MetaHuman named Jin. According to Nvidia, Jin answers player questions in natural language and responds in a way that fits the character's background story.
The demo's low latency stood out, but the article reports that details about the language model and the hardware running behind the scenes were not available. That leaves open how the demonstration's speed and behavior would translate to different games and development setups.
Convai founder and CEO Purnendu Mukherjee said the tools could help make AI non-playable characters available to more developers, while meeting expectations for latency and quality in a cost-efficient way. Those are goals for the technology; the demonstration alone does not establish how broadly studios will adopt it.
Developers can choose how to deploy models
Nvidia says ACE models vary in size, performance, and quality. Its ACE for Games foundry service is intended to help developers fine-tune models for their games, then deploy them through Nvidia DGX Cloud, GeForce RTX PCs, or on premises for real-time inference.
The company says the models are optimized for latency, which it describes as a critical requirement for immersive, responsive game interactions. The deployment choices indicate that studios may be able to fit the system to different production needs, though the source does not specify the hardware or costs a particular game would require.
One announced use is focused on animation
GSC Game World is among the first development studios named as using an ACE tool. Audio2Face is planned for S.T.A.L.K.E.R. 2 Heart of Chernobyl, showing a potential use for generated facial animation in a commercial game.
The source draws a boundary around what that means: AI-generated campfire stories from NPCs are not planned for the time being. Using an animation component does not necessarily mean a game will generate its characters' stories or dialogue dynamically.
For players, the promise of ACE is a character who can respond in speech and show matching facial movement. For developers, the challenge is making those responses quick, consistent with the character, and practical to deploy. Nvidia's announcement describes tools aimed at that challenge, while leaving the exact performance and implementation details dependent on the game.