sungikchoi_da34c0c41.jpg Sungik Choi 2024.02.01

[NeurIPS 2023] Trends on Score-based Generative Modeling

최근 Score-based Generative Modelling 혹은 Diffusion Model 연구의 유용한 특성과 뛰어난 성능이 주목받으면서 관련 연구 결과도 활발하게 발표되고 있습니다. NeurIPS 2023에서도 해당 주제를 다룬 논문들이 다양하게 발표 되었습니다. 이번 글을 통해 NeurIPS 2023에서 확인한 Score-based Generative Modeling 연구 트렌드를 살펴보고, 인상 깊었던 연구 논문을 소개하고자 합니다.


1. Diffusion Model Architecture

Diffusion Model의 Loss는 겉보기에는 간단한 L2-loss로 되어 있으나 이 안에는 다양한 Weight와 Hyperparameter들이 들어있습니다. 이러한 Hyperparameter 설정에 대한 내용을 전반적으로 다룬 엔비디아(NVIDIA)의 EDM[1]이 지난 NeurIPS 2022에서 Outstanding Paper로 지정되었습니다. 이번 NeurIPS 2023에서도 구글 딥마인드(Google DeepMind)의 “Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation”[2] 이 Oral Presentation으로 선정되었습니다. 해당 논문에서는 각기 다른 Noise Level의 Diffusion Model에서 Training Loss를 합치는 Weighting 설정에 대해 다룹니다.



만약 Weighting이 1이라면 이는 ELBO Term에 해당하게 되지만 대부분의 Diffusion Model들은 목적 (Perceptual Quality)에 따라 다른 Weighting을 쓰게 됩니다.

 

그림 1. 현재 Diffusion Model들에서 사용하는 서로 다른 Weighting[2]

 

본 연구에서는 Monotonic Weighting이 Gaussian Perturbation 형태의 Data Augmentation 됐을 때의 ELBO로 나타낼 수 있다는 것을 증명하고, EDM Objective의 Monotonic한 형태의 Weighting을 디자인합니다. 저자들은 다양한 Resolution의 ImageNet Dataset에서 해당 Weighting을 적용하여 기존 Scheduling 대비 일관적이고도 유의미한 성능 향상을 이루었습니다.

위 논문 외에도 Diffusion Model의 Base Architecture에 대한 U-Net Architecture에 대해 성능적으로 접근한 두 논문들[3,4] 또한 인상 깊었습니다.

 

▶ C. Williams et al. A Unified Framework for U-Net Design and Analysis, NeurIPS 2023[3] (Link)

Image 기반 Diffusion Model에서 빼놓을 수 없는 Base 구조는 U-Net Architecture입니다. 본 연구에서는 기존 Residual U-Net에서 Encoder Part를 Wavelet 등의 Basis Function으로 구성하여 학습하고자 하는 Parameter를 줄입니다. 해당 구성은 다양한 Basis를 사용할 수 있게 하고 Sphere 등의 Geometry 등에도 적용할 수 있는 General한 U-Net 구조를 만듭니다. 본 논문은 실험을 통해 같은 Hyperparameter 크기 대비 설계한 Multiresolution U-Net의 우수한 성능을 보입니다.


▶ Y. Li et al., SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds, NeurIPS 2023[4] (Link)

Text-to-Image Diffusion Model로 원하는 프롬프트를 만들 때 인터넷 서버에서 몇 분씩 기다리면서 돌려보신 경험 있으실 겁니다. Text-to-Image Model의 Sampling 속도가 빨라졌다고는 하나, 여전히 몇 분씩 걸리는 긴 Sampling 속도는 한계로 자리하고 있습니다. 해당 논문은 Large Text-to-Image Model의 U-Net Decoder 부분을 Data Distillation을 통해 Sampling 속도를 높이는 방법을 소개합니다. 또한 다양한 Training Strategy와 결합하여 On-device에서도 처리 가능한 효과적이고 빠른 Text-to-Image Model을 생성합니다.

 

2. Text-to-Image Diffusion Model : Downstream Task

CVPR 2022에 Latent Diffusion Model (LDM) 및 Stable Diffusion[5]이 소개된 이후로 Text-to-Image Diffusion Model을 활용하는 연구가 많이 이루어졌습니다. 이번 NeurIPS 2023에서도 기존의 Generation 뿐만 아니라 다양한 Downstream Task에 적용하는 논문들이 소개됐습니다. 그 중에서 제가 소개하고 싶은 논문은 구글 딥마인드에서 나온 “Text-to-Image Diffusion Models are Zero-Shot Classifiers”[6] 논문입니다.

 

그림 2. Zero-Shot Classification 과정[6]

 

본 논문은 각 Class에 관련된 Prompt를 Condition으로 주어 각 결과를 Guidance로 만들어냅니다. 그 이후 기존 모델과의 차이를 Score로 주어 여러 Timestep에서 합산하여 최소의 차이를 내는 Class를 Output으로 쓰는 형태로 Zero-Shot Classification을 진행합니다. 해당 방식은 기존 CLIP 등 모델의 Zero-Shot Classification보다 좋은 성능을 보였고 Classifier가 Overfitting되기 쉬운 Texture Bias에도 조금 더 Robust한 결과를 보였습니다. 비록 Diffusion Model 기반이라 Evaluation 시간이 많이 들기는 하나 충분히 기대되는 방향으로 보입니다.

위 논문 외에도 Goal-conditioned RL[7], Point-cloud Completion[8] 등의 논문들도 Text-to-Image Diffusion Model의 여러 Downstream Task로의 가능성을 확인할 수 있는 흥미로운 연구였습니다. 여러 Task와 Condition Modality를 한 Diffusion Model로 처리하는 연구도 있었습니다. Uni-ControlNet[9]은 Pose, Segmentation Map 등의 Local Condition을 Local Control Adaptor를 통해 Stable Diffusion과 호환 가능하게 설정하였고, Global Control Adaptor를 통해서 Text 토큰에 추가적으로 Prompting Condition을 추가로 설정하였습니다. 본 모델은 다양한 Modality를 Condition하여서도 기존 Baseline 대비 좋은 성능을 보였습니다.


▶ J. Gao et al., Can Pre-trained Text-to-Image Models Generate Visual Goals for Reinforcement Learning?, NeurIPS 2023[7] (Link)

Score-based Generative Model을 강화학습(Reinforcement Learning, RL)에 적용한 사례는 2022년 ICML에서 발표된 논문[10] 이후로 크게 증가했습니다. 하지만 대부분의 논문들은 (State, Action) Trajectory를 모델링하거나[11] Policy[12], 혹은 이 Distribution의 Hyperparameter[13]를 Modeling하는 경우가 많았으며 Architecture도 비교적 단순한 모델을 사용했습니다. 본 연구는 Pre-trained된 Latent Diffusion Model과 Language Instruction을 통해 강화학습 Task에서 도달하고자 하는 Goal State를 Generation하였고 이를 Reinforcement Learning Backbone Algorithm에 Positive Sample로 Augment 시킵니다. Text-to-Image Diffusion 모델의 다양한 생성 가능성을 RL로 확장한 인상적인 연구입니다.

LG AI연구원의 Advanced Machine Learning(AML) 랩에서는 Diffusion Model을 포함하여 다양한 생성 모델에 대한 이론적인 연구뿐만 아니라, 다양한 Downstream Task에 모델을 적용하는 가능성에 대한 연구도 폭넓게 진행하고 있습니다. 앞으로도 글로벌 연구 트렌드를 분석하고, 실용성 있는 기술 발전과 성과를 만들어낼 수 있도록 하겠습니다.


▶ NeurIPS 2023 Research Blog 시리즈가 궁금하다면? (Link)

참고

[1] T. Karras et al., Elucidating the Design Space of Diffusion-Based Generative Models, NeurIPS 2022

[2] D. P. Kingma and R. Gao., Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation, NeurIPS 2023

[3] C. Williams et al., A Unified Framework for U-Net Design and Analysis, NeurIPS 2023

[4] Y. Li et al., SnapFusion: Text-to-image Diffusion Model on Mobile Devices within Two Seconds, NeurIPS 2023

[5] R. Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models, CVPR 2022

[6] K. Clark and P. Jaini, Text-to-image Diffusion Models are Zero-Shot Classifiers, NeurIPS 2023

[7] J. Gao et al., Can Pre-trained Text-to-image Models Generate Visual Goals for Reinforcement Learning?, NeurIPS 2023

[8] Y. Kasten et al. Point-Cloud Completion with Pretrained Text-to-image diffusion models, NeurIPS 2023

[9] S. Zhao et al. Uni-ControlNet: All-in-one Control to Text-to-image Diffusion Models, NeurIPS 2023.

[10] M. Janner et al. Planning with Diffusion for Flexible Behavior Synthesis, ICML 2022

[11] A. Ajay et al. Is Conditional Generative Modeling all you need for Decision-Making?, ICLR 2023.

[12] Z. Wang et al. Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning, ICLR 2023.

[13] H. Chen et al. Score Regularized Policy Optimization through Diffusion Behavior, ICLR 2024