|
NVIDIA GTC is a global AI conference focusing on artificial intelligence, computer graphics, and machine learning. NVIDIA GTC 2022 was held online from September 19 to 22 where researchers and business leaders actively exchanged and communicated on topics such as large scale language models, digital twin, and human AI. Hyunmin Lee and Minsoo Lee from Vision Lab, LG AI Research, who took part in the online conference and talk about GTC 2022. |
In this post, we provide a summary of the GTC 2022 session pertaining to digital avatars.
The digital avatar market continues to grow annually. According to a report published by the market research firm Emergen Research in 2022, the market for digital avatars is expected to reach USD 527.76 billion by 2030. Digital avatars can be implemented as a variety of entities, ranging from robots to virtual humans, and are attracting a great deal of attention as they can be used in numerous industries, including education, gaming, and healthcare.
Digital avatars can be divided into three main categories: digital assistants, AI service agents, and virtual influencers. (1) We are likely most familiar with digital assistants. An example is instant chatbots powered by AI. (2) An AI service agent is a form of AI with an interactive animated avatar. By applying animations such as eye contact, facial expression changes, and body language, and through instant responses and domain knowledge, it comprehends context and provides appropriate recommendations. (3) A virtual influencer is a type of AI that applies technology related to voice, facial expression, and gestures to a 3D avatar representing a person, which is used in video games, on v-tube, etc.
Creating an interactive digital avatar requires 3D avatar creation, AI model training, and application development skills, making the overall process complex and time-consuming. Nvidia has therefore developed the Omniverse ACE (Avatar Cloud Engine) to facilitate the creation and deployment of avatars. ACE is a suite of cloud-based AI models and services that enable the creation and deployment of interactive, realistic avatars. We will first describe the AI technologies (Audio2Face and Text-to-Speech) included in ACE before introducing the applications created using ACE (Tokkio and Maxine).
Audio2Face [Demo Video] [Introductory Website]

Figure 1. Animating facial features by controlling emotions such as anger, happiness, and sadness ⓒNVIDIA Omniverse™
Audio2Face is an AI technology based on deep learning that generates facial animations with expressions from audio input. The key characteristics include real-time performance, retargeting to other characters via face-swap, and post-processing of emotion control shown in Figure 1. Most impressively, it can generate facial expressions by directly inferring emotions from the audio clip.
Text-to-Speech [Demo Video] [Introductory Website]

Figure 2. Creating voices by adjusting pitch, speed, pronunciation, volume, etc. ⓒ NVIDIA Riva
TTS is a technology that accepts text as input, generates a voice, and makes the avatar speak the generated voice. The objective is to create a natural human-sounding voice, as opposed to a machine sound. It has the ability to create a customized voice that adjusts pitch, speed, pronunciation, and volume through a 30-minute voice file.
Tokkio (Customer Service Kiosk) [Demo Video]

Figure 3. Tokkio, a customer-facing service employing multiple AI technologies ⓒNVIDIA Omniverse ACE
Tokkio combines multiple AI technologies to provide customers with a variety of services. Included are technologies such as computer vision, interactive AI, 3D animation, and recommendations, thereby allowing the AI to accept orders and recommend menu items as if it were a store clerk.
Maxine (Video Communication Framework) [Demo Video] [Introductory Website]

Figure 4. Eye position adjusted to face the camera during video conferencing sessions ⓒNVIDIA Maxine

Figure 5. Conversation in multiple languages with translation and Audio2Face technology ⓒNVIDIA Maxine
Maxine is an application that facilitates clear communication during video conference, customer service calls, and other interactions. Maxine utilizes speech recognition and translation technologies, including Audio2Face and Text-to-Speech, which were introduced previously. As depicted in Figure 4, the primary function is the adjustment of the positions of the eyes so that they face the camera, and as shown in Figure 5, automatic translation technology enables communication with people from around the world.
Conclusion
This post examined the various types of digital avatars and the required skills. Digital avatars are gradually gaining the ability to communicate with human-like facial expressions and voices as a result of the constant development of AI technology. We anticipate the future development of applications that integrate multiple AI technologies, such as Tokkio and Maxine, introduced by Nvidia.