CLIP(Contrastive Language-Image Pre-training)[9]은 인터넷 상에 존재하는 대규모의 image-text pair 데이터를 사전학습한 모델입니다. 다양한 이미지 분류 작업에 범용으로 사용될 수 있고, 특히 뛰어난 zero-shot 성능(학습 과정에 사용되지 않은 데이터세트를 분류하는 성능)을 보여줍니다. 예를 들어 CLIP은 사람이 레이블한 약 120만 장의 ImageNet[3] 데이터를 학습한 모델과 유사한 성능으로, 처음보는 ImageNet 이미지를 분류할 수 있습니다.
하지만 이런 모델을 학습시키기 위해서는 ImageNet의 400배에 해당하는 약 4억 개의 image-text pair 데이터와 수백 개의 GPU에 해당하는 연산 능력이 필요합니다. 때문에 시간과 비용이 제한된 상황에서는 학습이 매우 어려운 한계점이 있습니다.
본 연구에서는 이러한 한계를 보완하기 위해 CLIP과 같은 large-scale pre-trained multimodal model의 representation을 ResNet-18과 같은 소규모 target model로 transfer 시킬 수 있는 방법인 BeamCLIP을 소개합니다. 이를 위해서, 본 연구에서는 Cross-modal Similarity Matching(CSM)과 Context-based Prompt Augmentation(CPA)이라는 두 가지 방법을 제안합니다.
아래 Figure 1은 ResNet-18에서 ImageNet top-1 linear probe 성능 비교 결과를 보여줍니다. Linear probe 성능은 downstream task의 성능을 판가름하는 척도와 같습니다. 비전기반 Self-Supervised Learning(SSL)방법인 SimCLR[2]와 SwAV[1]는 각각 51.8%와 63.7%의 성능을 나타냈고, 본 연구에서 제안하는 BeamCLIP은 66.2%의 성능을 기록했습니다. 이는 BeamCLIP이 비전기반 SSL 방법들보다 우수한 representation을 학습할 수 있다는 점을 시사합니다.

개념적으로, BeamCLIP은 unsupervised representation transfer 방법으로 볼 수 있습니다. 어떤 unlabeled query image에 대해, student 모델이 teacher 모델의 vector representation을 regress하도록 학습하는 것입니다. Vector regression을 위해서, 기본적으로 L2 distance, 즉 cosine similarity를 사용할 수 있습니다. 이때 발생할 수 있는 topological ambiguity 문제를 완화하기 위해 Cross-modal Similarity Matching(CSM)을, lexical ambiguity 없이 text prompt를 적절하게 인코딩하기 위해 Context-based Prompt Augmentation(CPA)을 사용합니다.
Figure 2는 BeamCLIP의 구조와 방법을 개략적으로 보여줍니다. BeamCLIP은 parameter를 고정한 CLIP[9]을 teacher 모델로 사용해서 student 모델을 학습시킵니다. CLIP은 이미지 인코더와 텍스트 인코더로 구성되어 있습니다. 이 인코더들은 사전학습 되어 있고, 각 인코더의 representation은 cross-modal embedding space에서 well-aligned 되어 있습니다. 
Cross-modal Similarity Matching(CSM)
Student model이 teacher model의 vector representation을 regress하기 위해서, BeamCLIP은 기본적으로 instance similarity matching(ISM)을 사용합니다. (Figure 3의 수식 참고)
Figure 4에는 Cross-modal Similarity Matching(CSM)의 개념도가 있습니다. 어떤 query image가 주어지면, student 모델의 vector representation이 teacher 모델의 vector representation과 cosine similarity가 높아지도록 학습시킵니다. 하지만, instance similarity matching은 각도가 같고 방향이 다른 vector representation을 optimize하지 못한다는 topological ambiguity 문제가 있습니다. 이러한 topological ambiguity 문제를 완화하기 위해서, BeamCLIP은 text prompt의 vector representation과의 cosine similarity도 함께 optimize합니다.
예를 들어, teacher model의 embedding space에서 개 이미지는 ‘a photo of dog’이라는 text prompt와는 cosine similarity가 높고, ‘a photo of cat’이라는 text prompt와는 cosine similarity가 낮습니다. BeamCLIP은 이러한 teacher model의 cross-modal similarity를 student model이 유지하도록 학습시킵니다.
Cross-modal similarity matching을 좀 더 구체적으로 설명해 보겠습니다. 첫째로, teacher의 embedding space에서 이미지와 텍스트의 상대적인 관계를 알기 위한 방법은 다음과 같습니다. 어떤 query image가 주어지면, 주어진 text prompt에 대한 normalized cross-modal similarity를 측정합니다. (Figure 5의 수식 참고) 이 normalized cross-modal similarity를 이용해서, 주어진 text prompt 집합에 대한 cross-similarity distribution을 계산합니다. (Figure 6의 수식 참고)
둘째로, student 모델은 teacher 모델과 같은 cross-modal similarity distribution을 갖도록 학습합니다. 이때, cross-entropy를 이용해서, student와 teacher의 cross-modal similarity distribution을 맞춥니다. (Figure 7 참조)
Context-based Prompt Augmentation(CPA)
Figure 8는 context-based prompt augmentation(CPA)의 개념도를 보여줍니다. BeamCLIP은 보다 정교하게 teacher representation을 학습하기 위해서, text prompt의 vector representation을 anchor로 사용합니다. 하지만, text prompt에 존재할 수 있는 lexical ambiguity 때문에, 의미적으로 부정확한 vector representation이 생성될 수 있습니다.
예를 들어, Flowers102[8] 데이터세트에는 ‘snapdragon’이라는 class name이 있습니다. 이것만 봐서는 꽃 종류인지 동물 종류인지 모호합니다. 이런 경우, ‘a photo of snapdragon’이라는 text prompt의 vector representation은 ‘a photo of flower’보다 ‘a photo of dragon’과 cosine similarity가 높을 수 있습니다. 이러한 lexical ambiguity를 완화하기 위해 Context-based prompt augmentation(CPA)은 basic text prompt를 Wikipedia description이나 hierarchical label을 활용합니다.
Experiments
비전기반 Self-Supervised Learning (SSL)과의 비교

Figure 9은 ResNet-18에서 ImageNet-1K top-1 linear probe accuracy 성능을 보여줍니다. BeamCLIP은 100 epoch 학습 후 63.8%의 성능을 달성하면서, 비전기반 self-supervised learning 방법인 SimCLR[2](51.8%), BYOL[4](61.4%), SwAV[1](63.7%)보다 우수한 representation을 학습할 수 있음을 보여줍니다. 특히, ResNet-18과 같이 소규모 모델에서 이와 같은 성능을 냈다는 점에 주목할 필요가 있습니다. Self-supervised learning은 소규모 모델에서 잘 동작하지 않는다고 알려져 있기 때문입니다. BeamCLIP은 CLIP[9]과 같은 pre-trained multimodal model을 teacher 모델로 이용함으로써 ResNet-18과 같은 소규모 모델에서도 좋은 representation을 학습하는 것이 가능함을 보여줍니다.
Zero-shot 성능 비교

Figure 10은 ImageNet-1K[3] 데이터세트에 대한 zero-shot 성능을 비교한 결과입니다. ImageNet 데이터세트 상에서, BeamCLIP-RN50은 57.5%의 zero-shot 성능을 달성합니다. 이는 59.6%의 CLIP-RN50과 견줄 만한 zero-shot 성능입니다.
여기서 중요한 부분은 이러한 zero-shot 성능을 얻기 위해 CLIP[9]은 약 4억 개의 image-text pair 학습 데이터가 필요한 반면, BeamCLIP은 약 3%에 해당하는 데이터만 사용해도 된다는 점입니다. 약 1,200만 개의 unlabeled image를 (즉, ImageNet-21K에서 ImageNet-1K를 제외한 데이터세트) 학습 데이터로 사용해서, CLIP과 유사한 zero-shot 성능을 얻을 수 있다는 의미입니다. 참고로, 동일한 조건으로 약 1,200만 개의 image-text pair 학습 데이터를 사용할 경우, CLIP의 오픈 소스 버전인 OpenCLIP[6]은 단지 약 36.5%의 성능밖에 내지 못합니다.
기존 Knowledge Distillation(KD)과의 비교
Figure 11은 다양한 데이터세트에 대한 teacher와 student 모델의 성능 비교를 보여줍니다. KL divergence를 사용하는 기존 Knowledge Distillation[5]이나, contrastive loss를 사용하는 Contrastive Representation Distillation(CRD)[10]는 CLIP representation transfer에 효과적이지 않은 것으로 나타납니다. 반면에 Cross-Modal Similarity Matching(CMS)을 이용하는 BeamCLIP은 unlabeled query image를 사용해서 CLIP representation을 ResNet-50으로 효과적으로 transfer할 수 있음을 보여줍니다.
정성적 비교
Figure 12는 Flowers102[8] 데이터세트에서 CLIP-RN50와 BeamCLIP-RN50의 top-5 text-image retrieval 결과 비교를 보여줍니다. 주어진 text prompt와 매치되지 않는 이미지는 빨간색 사각형으로 표시했습니다. 보다시피 CLIP-RN50은 부정확한 결과가 상당히 많습니다. 반면에 BeamCLIP-RN50은 zero-shot 성능이 높은 CLIP-ViT-B/16의 우수한 representation을 ResNet-50으로 transfer해서 정확도가 크게 높아진 것을 확인할 수 있습니다. 특히 흥미로운 점은 ‘a photo of yellow rose’와 같은 unseen text prompt에 대해서도 높은 성능을 보인다는 점입니다.
Conclusion
본 연구에서는 CLIP[9]과 같은 pre-trained multimodal representation을 ResNet-18과 같은 소규모 모델로 효과적으로 transfer할 수 있는 방법인 BeamCLIP을 제안하고 있습니다. BeamCLIP은 image와 text prompt 사이의 cross-modal similarity를 유지시키는 cross-modal similarity matching (CSM)과 text prompt를 Wikipedia 설명으로 확장하는 context-based prompt augmentation(CPA)를 사용해, 효과적으로 CLIP representation을 transfer할 수 있음을 실험적으로 입증했습니다. BeamCLIP은 대규모 학습 데이터와 컴퓨팅 자원을 필요로 하는 large-scale multimodal model의 우수한 representation을 효율적으로 재활용함으로써, multimodal AI application 개발에 필요한 비용과 시간을 줄일 수 있을 것으로 기대하고 있습니다. 본 연구를 기반으로, LG AI연구원은 다양한 문제들을 보다 효과적이고 효율적으로 학습할 수 있는 기계학습 방법론 관련 연구를 이어나갈 계획입니다.
▶ Transferring Pre-trained Multimodal Representations with Cross-modal Similarity Matching (Link)