There is an old saying in Korea, “if the body is worth a thousand, eyes are worth nine hundred”. It stresses the importance of our eyes as most information is obtained through them. Computer vision is an AI technology in which machine recognizes objects and space through images or video data from camera sensors, in the same way humans interpret and understand objects and space through vision. Let us meet with Seung Hwan Kim, who leads LG AI Research's Vision Lab and hear about computer vision research and Vision Lab.
Introduction to Deep Learning, Eyes of the Computer
Lab leader Seung Hwan Kim is a computer vision expert who has researched image/video recognition and processing, and video cloud for 25 years. During the early stages of research, he carried out Handcraft Feature based research where people directly interpreted data characteristics mathematically, however, seeing the rapid growth of vision research through deep learning, he has been challenging for endless possibilites. After leading the AI technology organization at LG Uplus, lab leader Kim joined LG AI Research when it was first launched and has been leading Vision Lab since.
Q. In comparison to other AI sectors, what are some characteristics or differentiating factors of the vision field?
The image data that we use in the vision field causes displacement according to environmental factors. Such data includes both the background and the object being recognized, and sometimes changes in the background occur as well. For example. when recognizing a car, an image of a car in a tunnel and a car on the road during the day is recognized differently despite being the same car, as the colors and the background appear differently. Thus, it is difficult to obtain high performance in all conditions so we should develop and optimize a technology that is suitable for specific sectors.
In the vision field, there are diverse fields of research, data and applications. Research based on basic task can be divided into image creation and conversion, object/face/action/emotion recognition, object detection and tracking, and data that needs analysis varies from 2D/3D image and video contents to biometric information, face and document. Fields of application also vary from autonomous vehicles, media contents analysis, to medical, industrial AI. We may not be able to perfectly apply this in all fields but with the introduction of deep learning, we have been able to respond to some extent without having specialized knowledge in domain.
Q. The research infrastructure must also be very important considering the large data capacity used in research.
That's right. The size of image and video data is very large so it takes up a lot of memory space and we need a considerable amount of GPU resources to operate a single model. In this respect, LG AI Research allows researchers to experiment freely as it is equipped with the suitable research environment. We built our research infrastructure relatively recently than existing AI organizations so we are able to research in the latest GPU environment.
Q. Which technology in vision research is in the center of attention?
General AI is a hot topic in the AI field. OpenAI unveiled General AI, GPT3 which performs at a level similar to humans by using enormous data and infrastructure incomparable to existing models in the natural language processing field, and this trend has recently shifted to the vision field. In January of this year, Open AI unveiled two new models 'CLIP(Contrastive Language-Image Pre-training)' and 'DALL-E' which combines computer vision and natural language processing technology. CLIP uses large scale models to solve image classification problems and DALL-E generates images and texts as results when text is input. Thus, a lot of research on Multi-modal representation is being conducted to figure out how to combine and express different elements such as texts and images.
I believe these two models showed both potential and limitations. Unlike the language field, vision field is at the early stages of using large-scale models so we still need more work on enhancing the performance, however it also shows possibility of creating strong pre-trained models that could solve diverse vision task problems when based on transformers.
Q. What are some unsolved technical challenges in the vision field?
Vision research has developed centering around specific technologies such as classification, detection, segmentation and video generation, and some technologies perform at the level of humans in an ideal environment. However, there is a limit to grasping the whole image or understanding and reasoning using visual information. For instance, when we see a photo of a cat with a person, we can derive that the person is the cat's owner, however the current vision technology is not at the level of understanding the details of the relationships between objects. So these are some of the challenges that are yet solved.
A Challenge that Goes Beyond the Limits of Human Vision
Last year, LG AI Research won first place at the CVPR 2020 Continual Learning Computer Vision Challenge. It was a remarkable achievement made through a joint research with the University of Toronto. CVPR is a world reknowned international conference in the computer vision field, and taking first place in this competition meant being recognized for our technological capabilities after competing with global companies that had started AI research before LG AI Research. In addition, LG AI research published papers on Continual Learning and Explainable AI(XAI) research in world's prestigious AI journals.

LG's CVPR 2020 Challenge Win Authentication Certificate
Q. Can you tell us about the researches being conducted at Vision Lab?
We are doing researches largely in the fields Visual Analytics, Visual Reasoning and Visual Generation. In the Visual Analytics field, we use the latest machine learning algorithms such as AutoEncoder, Contrastive Learning, Active Learning, Continual Learning, Explainable AI to solve problems on Classification, Detection and Recognition. This technology is used for vision inspection, deep document understanding and sports action recognition.
In the Visual Reasoning sector, we are conducting research on understading the video scenes through Graph Neural Network, Transformer based Text-Image Pair. In the Visual Generation field, we use technologies such as AutoEncoder, GAN, Transformer to solve problems in relation to Data Augmentation, 3D Reconstruction & Synthesis, Image-to-Image Translation and we are working to apply this technology to 3D virtual avatars, digital signages, and visual chatbots.
Q. Which project is the lab most focusing on?
We are currently concentrating on researching four sectors; vision inspection, deep document understanding, 3D Reconstruction & Synthesis and large scale vision model. Vision inspection automates quality inspection required in the manufacturing process using computer vision. In the past, the amount of defective data was small even when using AI, so it was difficult for machines to learn and there was a problem of having to label each data individually. We are solving this problem by having the models learn the correct data and identify as defective data when encountered with otherwise. We are currently co-working with LG Innotek and LG Electronics Production Engineering Research Institute and plan on applying this into our production sites within this year.
Deep document understanding is a crucial project for LG group to complete its digital transformation journey. We have been proceeding with digitalizing the documents of the chemicals subsidiaries. The current vision technology has a high recognition rate on printed text, however it is still difficult to recognize and structuralize graphs and charts. We are focusing our research on structuralizing chemical formulas and charts.
Recently, virtual influencers like Lil Miquela has been in the spotlight. We are also conducting research on 3D facial animation where we develop lip movement that matches the lip sync of virtual humans as well as facial expressions. Lastly, we also started Text-to-Image Generation, Scene Understanding research using large scale vision models where image is created when text is input.
Q. Were there any hardships while conducting the research or in the application process?
Frankly speaking, the visual results of vision research turn out clear, but there are many cases when the target is unclear when we try to apply it on-site. Thus, this makes it difficult for commercialization. Although there is a wide range of fields of applications, most challenges our subsidiaries face is in relation to numerical or language based data so there are only few areas where vision technology can be applied to.
In addition, vision research requires leadership and creativity. Just as people can interpret the same thing they see differently, it is difficult to digitally interpret the analogue information obtained from camera sensors. There are no given rules in video so there are parts that require creative ways to solve a problem. So for those who want to put their goofy creative ideas into action, visual research will be more fun to research.
I want to create the loudest organization of this research institute
Vision Lab consists of 12 members including lab leader Seung Hwan Kim. 8 members fall into the age range in their late 20s to early 30s. The average age range is pretty young even within LG AI Research which is known to be a 'young' organization. Owing to open recruitment where anyone competent and fit for LG AI Research is welcome to join, Vision Lab was able to attract talents with diverse backgrounds including foreign researchers to researchers who are carrying out their studies in graduate school while conducting research at the lab.
Q. What is the vision of Vision Lab?
The vision of our lab is 'Beyond Human Eyes, Real to Virtual', which means to challenge the limits of human vision in virtual space beyond reality. Vision is all about imitating human eyes. We are aiming to create AI that exceeds the level of human vision, and ultimately solve LG groups' business challenges and strengthen future business capabilities. We mainly use actual data but we plan on expanding our research scope to virtual spaces through avatars.
Q. What do you emphasize most to your lab members?
I often tell our members that there is nothing more important than to verify the results with their own 'eyes'. In the end, research is all about interpretation and debugging the results. Instead of relying on the numerical results, it is important to analyze how the results came out, and determine areas of improvement once this analysis is done properly. I believe depending on how creatively these areas are worked out makes an outstanding researcher.
In vision research, it is possible to verify the outcomes with our own eyes. It is crucial to take time and effort to check over the results until you find areas of improvement. You must see it with your own eyes and I believe gaining experience working out the improvements in a creative way becomes your competence.
Q. How are you planning on leading Vision Lab in the future?
I want our lab to be an organization where we talk a lot. A lot of the times, ideas come out in small talks rather than in formal meetings or discussions. I want to create an atmosphere where our members can focus not only on their research, but be able to meet with each other over coffee and talk comfortably about their daily lives and also exchange ideas on their research.
All members came from different backgrounds so despite being awkward at first, it has been getting much better now. I hope unique and brilliant talents join us so our lab becomes more noisy and energetic.
Q. What goals do you want to achieve at Vision Lab?
Vision is a common sector in the world. Natural language processing or voice sector is specialized in Korean so it may be difficult to attract global talents, however there are no entry barriers in the vision field. I want to make our lab a world-class vision research organization with diverse global talents. And I hope all of our members grow into top global talents. Furthermore, although it may be a bit late, I plan on creating a large scale vision model that can be presented on the global stage.
Personally as a researcher, it was infortunate that vision technology was being used in a limited area. I want to expand the application scope of vision technology by creating a model that can be universally used rather than one that is specialized for each task.
MINI INTERVIEW
Vision Lab Hu Yi(left), Sungha Choi(right)
Q. Can you tell us about your current work at Vision Lab?
Sungha: After starting my career at LG Electronics CTO division and working in the automobile R&D field for 10 years, I joined LG AI Research last November. I am currently researching Anomaly Detection which detects anomalies based on Constructive Learning to solve vision inspection problems.
Hu Yi: After graduating college in China, I have been researching computer vision while pursuing M.S and Ph.D studies in Korea. At LG AI Reserach, I am researching audio based 3D facial animation model. Humans' emotion is intangible and at times we are unable to accurately sense it which makes modeling very challenging. Our goal is to create digital avatars that are capable of expressing emotions like humans.
Q. What appealed to you about vision research?
Sungha:The fact that image is visiually more intuitive than language when solving a problem was appealing. Among vision technologies, I am particularly interested in deep learning based technology that can be applied in the automotive field, so during my master's studies, I conducted Semantic Segmentation, Domain Generalization, and Image Translation research based on Gan that is mainly applied in automobiles.
Hu Yi: Vision is the most crucial way for humans to perceive and understand the world. If a computer is able to properly interpret the world through the lens, it would be helpful in carrying out diverse automation tasks such as object detection and autonomous navigation. Also, visual contents like digital avatars created through generative models will interact naturally with humans, and thus bring revolution in the future AR/VR industry. I believe vision AI research will have meaningful impact and value.
Q. What kind of organization is Vision Lab? What kind of lab leader is Seunghwan Kim?
Sungha: Our Lab Leader Seunghwan Kim respects all members' opinions. As the project grows, it is crucial to achieve consensus with members through discussions and our lab leader creates an atmosphere for all members to speak comfortably. Through free discussions and reaching consensus, I believe Vision Lab will have a great impact on not only LG AI Research, but across all LG subsidiaries.
Hu Yi: Our lab leader communicates effectively with us based on his broad knowledge and in-depth expertise. With his sense of humor, he encourages our members for creative ideas and also supports us by giving us authority when we carry out our duties. The flexible, balanced workload, friendly and talented colleagues, and our leader's outstanding leadership motivates me to work harder in Vision Lab.
Q. What goals do you want to achieve at Vision Lab?
Sungha: I want to innovate the manufacturing process across all LG subsidiaries by succeeding the current project that I'm working on in the Anomaly Detection field. Furthermore, I would like Vision Lab and LG AI Research to secure world-class capabilities in specific fields of AI, and ultimately contribute to helping LG affiliates' businesses and generating new business opportunities.
Hu Yi: I am researching more universal ways to understand video and images at the level of humans. My long term goal is to develop a general visual representation model that compactly encodes information with multi-levels of semantics, structural relationship and layered layout.